# Table of Contents - [Unsloth Docs | Unsloth Documentation](#unsloth-docs-unsloth-documentation) - [Documentation Unsloth | Unsloth Documentation](#documentation-unsloth-unsloth-documentation) - [Fine-tuning pour débutants | Unsloth Documentation](#fine-tuning-pour-d-butants-unsloth-documentation) - [Hackathon de Reinforcement Learning IA AMD avec Unsloth | Unsloth Documentation](#hackathon-de-reinforcement-learning-ia-amd-avec-unsloth-unsloth-documentation) - [Installation via Conda | Unsloth Documentation](#installation-via-conda-unsloth-documentation) - [Mettre à jour Unsloth | Unsloth Documentation](#mettre-jour-unsloth-unsloth-documentation) - [Google Colab | Unsloth Documentation](#google-colab-unsloth-documentation) - [FAQ + Le fine-tuning est-il fait pour moi ? | Unsloth Documentation](#faq-le-fine-tuning-est-il-fait-pour-moi-unsloth-documentation) - [Installation d'Unsloth | Unsloth Documentation](#installation-d-unsloth-unsloth-documentation) - [Installer Unsloth sur MacOS | Unsloth Documentation](#installer-unsloth-sur-macos-unsloth-documentation) - [Exigences Unsloth | Unsloth Documentation](#exigences-unsloth-unsloth-documentation) - [Comment fine-tuner des LLM dans VS Code avec Unsloth et les GPU Colab | Unsloth Documentation](#comment-fine-tuner-des-llm-dans-vs-code-avec-unsloth-et-les-gpu-colab-unsloth-documentation) - [Installer Unsloth via Docker | Unsloth Documentation](#installer-unsloth-via-docker-unsloth-documentation) - [Hugging Face Hub, XET debugging | Unsloth Documentation](#hugging-face-hub-xet-debugging-unsloth-documentation) - [Continued Pretraining | Unsloth Documentation](#continued-pretraining-unsloth-documentation) - [Unsloth Environment Flags | Unsloth Documentation](#unsloth-environment-flags-unsloth-documentation) - [Multi-GPU Fine-tuning with Unsloth | Unsloth Documentation](#multi-gpu-fine-tuning-with-unsloth-unsloth-documentation) - [Unsloth Benchmarks | Unsloth Documentation](#unsloth-benchmarks-unsloth-documentation) - [Finetuning from Last Checkpoint | Unsloth Documentation](#finetuning-from-last-checkpoint-unsloth-documentation) - [GPU Mode - Reinforcement Learning Mini Conference 2026 | Unsloth Documentation](#gpu-mode-reinforcement-learning-mini-conference-2026-unsloth-documentation) - [Quantization-Aware Training (QAT) | Unsloth Documentation](#quantization-aware-training-qat-unsloth-documentation) - [How to Connect OpenRouter to Unsloth: API Key & Model Setup | Unsloth Documentation](#how-to-connect-openrouter-to-unsloth-api-key-model-setup-unsloth-documentation) - [500K Context Length Fine-tuning | Unsloth Documentation](#500k-context-length-fine-tuning-unsloth-documentation) - [How to Run Diffusion Image GGUFs in ComfyUI | Unsloth Documentation](#how-to-run-diffusion-image-ggufs-in-comfyui-unsloth-documentation) - [Fine-tuning Embedding Models with Unsloth Guide | Unsloth Documentation](#fine-tuning-embedding-models-with-unsloth-guide-unsloth-documentation) - [How to Fine-tune LLMs with Unsloth & Docker | Unsloth Documentation](#how-to-fine-tune-llms-with-unsloth-docker-unsloth-documentation) - [Connect Anthropic to Unsloth: Run Claude Models in Local Chat | Unsloth Documentation](#connect-anthropic-to-unsloth-run-claude-models-in-local-chat-unsloth-documentation) - [Multi-GPU Fine-tuning with Distributed Data Parallel (DDP) | Unsloth Documentation](#multi-gpu-fine-tuning-with-distributed-data-parallel-ddp-unsloth-documentation) - [Fine-tuning LLMs with NVIDIA DGX Spark and Unsloth | Unsloth Documentation](#fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth-unsloth-documentation) - [Text-to-Speech (TTS) Fine-tuning Guide | Unsloth Documentation](#text-to-speech-tts-fine-tuning-guide-unsloth-documentation) - [Connect llama.cpp to Unsloth: Run GGUFs with llama-server | Unsloth Documentation](#connect-llama-cpp-to-unsloth-run-ggufs-with-llama-server-unsloth-documentation) - [Connect vLLM to Unsloth for Local Chat Inference | Unsloth Documentation](#connect-vllm-to-unsloth-for-local-chat-inference-unsloth-documentation) - [How to Run Local AI Models with OpenClaw | Unsloth Documentation](#how-to-run-local-ai-models-with-openclaw-unsloth-documentation) - [Vision Fine-tuning | Unsloth Documentation](#vision-fine-tuning-unsloth-documentation) - [Fine-Tuning LLMs on NVIDIA DGX Station with Unsloth | Unsloth Documentation](#fine-tuning-llms-on-nvidia-dgx-station-with-unsloth-unsloth-documentation) - [Troubleshooting & FAQs | Unsloth Documentation](#troubleshooting-faqs-unsloth-documentation) - [Chat Templates | Unsloth Documentation](#chat-templates-unsloth-documentation) - [How to Run Local AI Models with Hermes Agent | Unsloth Documentation](#how-to-run-local-ai-models-with-hermes-agent-unsloth-documentation) - [Connect OpenAI to Unsloth: Run GPT Models in Local Chat | Unsloth Documentation](#connect-openai-to-unsloth-run-gpt-models-in-local-chat-unsloth-documentation) - [Fine-tuning LLMs with Blackwell, RTX 50 series & Unsloth | Unsloth Documentation](#fine-tuning-llms-with-blackwell-rtx-50-series-unsloth-unsloth-documentation) - [3x Faster LLM Training with Unsloth Kernels + Packing | Unsloth Documentation](#3x-faster-llm-training-with-unsloth-kernels-packing-unsloth-documentation) - [How to Run Local AI Models with OpenCode | Unsloth Documentation](#how-to-run-local-ai-models-with-opencode-unsloth-documentation) - [How to Connect Ollama to Unsloth | Unsloth Documentation](#how-to-connect-ollama-to-unsloth-unsloth-documentation) - [Run Coding Agents with Local LLMs using Unsloth Start | Unsloth Documentation](#run-coding-agents-with-local-llms-using-unsloth-start-unsloth-documentation) - [Connect Curl & HTTP to Unsloth | Unsloth Documentation](#connect-curl-http-to-unsloth-unsloth-documentation) - [Connect API Providers & Model Servers to Unsloth | Unsloth Documentation](#connect-api-providers-model-servers-to-unsloth-unsloth-documentation) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Unknown](#unknown) - [Connecter Anthropic à Unsloth : exécuter des modèles Claude dans le chat local | Unsloth Documentation](#connecter-anthropic-unsloth-ex-cuter-des-mod-les-claude-dans-le-chat-local-unsloth-documentation) - [Comment connecter OpenRouter à Unsloth : clé API et configuration du modèle | Unsloth Documentation](#comment-connecter-openrouter-unsloth-cl-api-et-configuration-du-mod-le-unsloth-documentation) - [Inférence Unsloth | Unsloth Documentation](#inf-rence-unsloth-unsloth-documentation) - [Connecter vLLM à Unsloth pour l'inférence de chat local | Unsloth Documentation](#connecter-vllm-unsloth-pour-l-inf-rence-de-chat-local-unsloth-documentation) - [Dépannage de l'inférence | Unsloth Documentation](#d-pannage-de-l-inf-rence-unsloth-documentation) - [Connecter OpenAI à Unsloth : exécuter des modèles GPT dans le chat local | Unsloth Documentation](#connecter-openai-unsloth-ex-cuter-des-mod-les-gpt-dans-le-chat-local-unsloth-documentation) - [Guide de déploiement du point de terminaison llama-server et OpenAI | Unsloth Documentation](#guide-de-d-ploiement-du-point-de-terminaison-llama-server-et-openai-unsloth-documentation) - [Enregistrer des modèles pour Ollama | Unsloth Documentation](#enregistrer-des-mod-les-pour-ollama-unsloth-documentation) - [Benchmarks GGUF Qwen3.5 | Unsloth Documentation](#benchmarks-gguf-qwen3-5-unsloth-documentation) --- # Unsloth Docs | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/get-started/readme.md) . Unsloth lets you run and train AI models on your own local hardware via an open-source UI. Our docs will guide you through running & training your own LLM locally. [Get started](https://unsloth.ai/docs/new/studio) [Our GitHub](https://github.com/unslothai/unsloth) [](https://unsloth.ai/docs/basics/amd) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FvyECRqXbIeD52Q4t9a04%252FAMD%2520promo%2520pic%25201920.png%3Falt%3Dmedia%26token%3D36563f7c-91ea-4b0f-a4b4-8c721e574808&width=490&dpr=3&quality=100&sign=df4d6005&sv=2) **Unsloth for AMD!** You can now run & train models on AMD. [](https://unsloth.ai/docs/models/glm-5.2) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FmbYXj0v0p5zbeESPDYUr%252Fglm52.png%3Falt%3Dmedia%26token%3D0cb4ae58-d249-403a-9fb5-72e9e93f8da9&width=490&dpr=3&quality=100&sign=a166d80d&sv=2) **GLM-5.2** Run the strongest open model locally. [](https://unsloth.ai/docs/integrations/unsloth-start) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FVB2vdala9hr9lGRgsf0B%252Funsloth%2520start%2520logo.png%3Falt%3Dmedia%26token%3Dadd879d3-e53b-4480-ba7f-8d4759f60123&width=490&dpr=3&quality=100&sign=54ed7bb2&sv=2) **Unsloth Start** Connect to your agent to any local LLM. [](https://unsloth.ai/docs/models/deepseek-v4) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FYCbdIeMsEQmDJoyii4tl%252Fdeepseek%2520v4%2520logo.png%3Falt%3Dmedia%26token%3D0d1ec333-bcbd-4ecb-8891-5393a1c5bb0a&width=490&dpr=3&quality=100&sign=6a6d3ec9&sv=2) **DeepSeek-V4** Run the new 284B Flash model. [](https://unsloth.ai/docs/basics/nvfp4) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FJdza44RRT0sgZcx9B67a%252Fdynamic%2520unsloth%2520nvfp4.png%3Falt%3Dmedia%26token%3D25d20d9d-bef4-43b5-b680-4ce1a37b4bd1&width=490&dpr=3&quality=100&sign=2d7c134&sv=2) **Dynamic NVFP4** Run models 2x faster on your Blackwell GPU. [](https://unsloth.ai/docs/new/studio) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FstfdTMsoBMmsbQsgQ1Ma%252Flandscape%2520clip%2520gemma.gif%3Falt%3Dmedia%26token%3Deec5f2f7-b97a-4c1c-ad01-5a041c3e4013&width=490&dpr=3&quality=100&sign=e4b21b2d&sv=2) **Introducing Unsloth Studio** New open, no-code UI to train and run LLMs. [Complete LLM Directory](https://unsloth.ai/docs/models/tutorials) [🧬Fine-tuning Guide](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide) [🔮Models](https://unsloth.ai/docs/get-started/unsloth-model-catalog) [Unsloth API](https://unsloth.ai/docs/basics/api) ### [](https://unsloth.ai/docs#quickstart) ⚡ Quickstart Unsloth supports MacOS, Linux, [Windows](https://unsloth.ai/docs/get-started/install/windows-installation) , [NVIDIA](https://unsloth.ai/docs/get-started/install/pip-install) , [AMD](https://unsloth.ai/docs/get-started/install/amd) , Intel and CPU setups. See: [Unsloth Requirements](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements) . Use the same commands to update: **MacOS, Linux, WSL:** Copy curl -fsSL https://unsloth.ai/install.sh | sh **Windows PowerShell:** Copy irm https://unsloth.ai/install.ps1 | iex ### [](https://unsloth.ai/docs#unsloth-start) 👾 Unsloth Start [Unsloth Start](https://unsloth.ai/docs/integrations/unsloth-start) lets you connect [Claude Code](https://unsloth.ai/docs/basics/claude-code) , [Codex](https://unsloth.ai/docs/basics/codex) and other agents to local models via the `unsloth start` command. Start Unsloth, load a model, open your project folder, and then run: Copy unsloth start claude Replace `claude` with any agent below: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpXE6kCHjh8qOEaggf94M%252FScreenshot_20260718_122426.png%3Falt%3Dmedia%26token%3Da59e4c8c-efdb-451b-b1f8-621955564f6d&width=768&dpr=3&quality=100&sign=4c664a42&sv=2) Claude Code running with Qwen3.5 locally. Agent Command Claude Code `unsloth start claude` OpenAI Codex `unsloth start codex` Hermes Agent `unsloth start hermes` OpenClaw `unsloth start openclaw` OpenCode `unsloth start opencode` ### [](https://unsloth.ai/docs#why-unsloth) 🦥 Why Unsloth? * We directly collab with teams behind [gpt-oss](https://docs.unsloth.ai/new/gpt-oss-how-to-run-and-fine-tune#unsloth-fixes-for-gpt-oss) , [Qwen3](https://www.reddit.com/r/LocalLLaMA/comments/1kaodxu/qwen3_unsloth_dynamic_ggufs_128k_context_bug_fixes/) , [Llama 4](https://github.com/ggml-org/llama.cpp/pull/12889) , [Mistral](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/discussions/18) , [Gemma 1-3](https://news.ycombinator.com/item?id=39671146) and [Phi-4](https://unsloth.ai/blog/phi4) , where we’ve **fixed critical bugs** that greatly improved model accuracy. Andrej Karpathy for example has [praised our work](https://x.com/karpathy/status/1765473722985771335) . * Unsloth streamlines local training, inference, data, and deployment * Unsloth supports inference and training for 500+ models: [vision](https://unsloth.ai/docs/basics/vision-fine-tuning) , [TTS](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning) , [embedding](https://unsloth.ai/docs/basics/embedding-finetuning) , [RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide) ### [](https://unsloth.ai/docs#features) ⭐ Features Unsloth lets you run and train models for text, [audio](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning) , [embedding](https://unsloth.ai/docs/new/embedding-finetuning) , [vision](https://unsloth.ai/docs/basics/vision-fine-tuning) and more. Unsloth provides many key features for both inference and training: #### [](https://unsloth.ai/docs#inference) Inference * [Self-healing tool calling](https://unsloth.ai/docs/new/studio/chat#auto-healing-tool-calling) / web search and use [Unsloth as an API](https://unsloth.ai/docs/basics/api) . * Connect your local models to any agent: [Claude Code](https://unsloth.ai/docs/basics/claude-code) , [Codex](https://unsloth.ai/docs/basics/codex) , [Hermes](https://unsloth.ai/docs/integrations/hermes-agent) and more. * Search + download + run any model like GGUFs, LoRA adapters, safetensors. * [Auto inference parameter](https://unsloth.ai/docs/new/studio/chat#auto-parameter-tuning) tuning and edit chat templates. * [Export or save](https://unsloth.ai/docs/new/studio/export) your model to GGUF, 16-bit safetensor etc. * [Compare outputs](https://unsloth.ai/docs/new/studio/chat#model-arena) with two different model side by side. #### [](https://unsloth.ai/docs#training) Training * Train and [RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide) 500+ models ~2x faster with ~70% less VRAM (no accuracy loss) * Supports full fine-tuning, pre-training, 4-bit, 16-bit and FP8 training. * [Auto-create datasets](https://unsloth.ai/docs/new/studio/data-recipe) from PDF, CSV, DOCX files. Edit data in a visual node workflow. * Observability: Monitor training live, track loss, GPU usage, customize graphs * Most efficient [**reinforcement learning**](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide) library, using 80% less VRAM for GRPO, [FP8](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning) etc. * [Multi-GPU](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth) works but a much better version is coming! ### [](https://unsloth.ai/docs#what-is-fine-tuning-and-rl-why) What is Fine-tuning and RL? Why? [**Fine-tuning** an LLM](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide) customizes its behavior, enhances domain knowledge, and optimizes performance for specific tasks. By fine-tuning a pre-trained model (e.g. Llama-3.1-8B) on a dataset, you can: * **Update Knowledge**: Introduce new domain-specific information. * **Customize Behavior**: Adjust the model’s tone, personality, or response style. * **Optimize for Tasks**: Improve accuracy and relevance for specific use cases. [**Reinforcement Learning (RL)**](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide) is where an "agent" learns to make decisions by interacting with an environment and receiving **feedback** in the form of **rewards** or **penalties**. * **Action:** What the model generates (e.g. a sentence). * **Reward:** A signal indicating how good or bad the model's action was (e.g. did the response follow instructions? was it helpful?). * **Environment:** The scenario or task the model is working on (e.g. answering a user’s question). **Example fine-tuning or RL use-cases**: * Enables LLMs to predict if a headline impacts a company positively or negatively. * Can use historical customer interactions for more accurate and custom responses. * Fine-tune LLM on legal texts for contract analysis, case law research, and compliance. You can think of a fine-tuned model as a specialized agent designed to do specific tasks more effectively and efficiently. **Fine-tuning can replicate all of RAG's capabilities**, but not vice versa. [Unsloth Updates](https://unsloth.ai/docs/new/changelog) [🖥️Inference & Deployment](https://unsloth.ai/docs/basics/inference-and-deployment) [Unsloth Start](https://unsloth.ai/docs/integrations/unsloth-start) [🦥Dynamic 2.0 GGUFs](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-134302f2507d4313b9575917c9a43b0a0028856c%252Flarge%2520sloth%2520wave.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=4e1b22de&sv=2) [NextModels](https://unsloth.ai/docs/get-started/unsloth-model-catalog) Last updated 2 days ago Was this helpful? --- # Documentation Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/readme.md) . Unsloth vous permet d’exécuter et d’entraîner des modèles d’IA sur votre propre matériel local via une interface open source. Notre documentation vous guidera pour exécuter et entraîner votre propre LLM en local. [Commencer](https://unsloth.ai/docs/fr/nouveau/studio) [Notre GitHub](https://github.com/unslothai/unsloth) [](https://unsloth.ai/docs/fr/notions-de-base/amd) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FvyECRqXbIeD52Q4t9a04%252FAMD%2520promo%2520pic%25201920.png%3Falt%3Dmedia%26token%3D36563f7c-91ea-4b0f-a4b4-8c721e574808&width=490&dpr=3&quality=100&sign=773bcf25&sv=2) **Unsloth pour AMD !** Vous pouvez désormais exécuter et entraîner des modèles sur AMD. [](https://unsloth.ai/docs/fr/modeles/glm-5.2) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FmbYXj0v0p5zbeESPDYUr%252Fglm52.png%3Falt%3Dmedia%26token%3D0cb4ae58-d249-403a-9fb5-72e9e93f8da9&width=490&dpr=3&quality=100&sign=143fb6ad&sv=2) **GLM-5.2** Exécutez localement le modèle ouvert le plus puissant. [](https://unsloth.ai/docs/fr/integrations/unsloth-start) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FVB2vdala9hr9lGRgsf0B%252Funsloth%2520start%2520logo.png%3Falt%3Dmedia%26token%3Dadd879d3-e53b-4480-ba7f-8d4759f60123&width=490&dpr=3&quality=100&sign=3fd2a812&sv=2) **Unsloth Start** Connectez votre agent à n’importe quel LLM local. [](https://unsloth.ai/docs/fr/modeles/deepseek-v4) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FYCbdIeMsEQmDJoyii4tl%252Fdeepseek%2520v4%2520logo.png%3Falt%3Dmedia%26token%3D0d1ec333-bcbd-4ecb-8891-5393a1c5bb0a&width=490&dpr=3&quality=100&sign=63947d69&sv=2) **DeepSeek-V4** Exécutez le nouveau modèle Flash de 284B. [](https://unsloth.ai/docs/fr/notions-de-base/nvfp4) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FJdza44RRT0sgZcx9B67a%252Fdynamic%2520unsloth%2520nvfp4.png%3Falt%3Dmedia%26token%3D25d20d9d-bef4-43b5-b680-4ce1a37b4bd1&width=490&dpr=3&quality=100&sign=e168f854&sv=2) **Dynamic NVFP4** Exécutez des modèles 2x plus vite sur votre GPU Blackwell. [](https://unsloth.ai/docs/fr/nouveau/studio) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FstfdTMsoBMmsbQsgQ1Ma%252Flandscape%2520clip%2520gemma.gif%3Falt%3Dmedia%26token%3Deec5f2f7-b97a-4c1c-ad01-5a041c3e4013&width=490&dpr=3&quality=100&sign=d7fcf80d&sv=2) **Présentation d’Unsloth Studio** Nouvelle interface ouverte, sans code, pour entraîner et exécuter des LLM. [Complete LLM Directory](https://unsloth.ai/docs/fr/modeles/tutorials) [🧬Fine-tuning Guide](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) [🔮Models](https://unsloth.ai/docs/fr/commencer/unsloth-model-catalog) [Unsloth API](https://unsloth.ai/docs/fr/notions-de-base/api) ### [](https://unsloth.ai/docs/fr#demarrage-rapide) ⚡ Démarrage rapide Unsloth prend en charge MacOS, Linux, [Windows](https://unsloth.ai/docs/fr/commencer/install/windows-installation) , [NVIDIA](https://unsloth.ai/docs/fr/commencer/install/pip-install) , [AMD](https://unsloth.ai/docs/fr/commencer/install/amd) , les configurations Intel et CPU. Voir : [Exigences Unsloth](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) . Utilisez les mêmes commandes pour mettre à jour : **MacOS, Linux, WSL :** Copier curl -fsSL https://unsloth.ai/install.sh | sh **PowerShell sous Windows :** Copier irm https://unsloth.ai/install.ps1 | iex ### [](https://unsloth.ai/docs/fr#unsloth-start) 👾 Unsloth Start [Unsloth Start](https://unsloth.ai/docs/fr/integrations/unsloth-start) vous permet de connecter [Claude Code](https://unsloth.ai/docs/fr/notions-de-base/claude-code) , [Codex](https://unsloth.ai/docs/fr/notions-de-base/codex) et d’autres agents à des modèles locaux via la `unsloth start` commande. Démarrez Unsloth, chargez un modèle, ouvrez votre dossier de projet, puis exécutez : Remplacez `claude` par n’importe quel agent ci-dessous : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpXE6kCHjh8qOEaggf94M%252FScreenshot_20260718_122426.png%3Falt%3Dmedia%26token%3Da59e4c8c-efdb-451b-b1f8-621955564f6d&width=768&dpr=3&quality=100&sign=ead4ace2&sv=2) Claude Code s’exécutant localement avec Qwen3.5. Agent Commande Claude Code `unsloth start claude` OpenAI Codex `unsloth start codex` Agent Hermes `unsloth start hermes` OpenClaw `unsloth start openclaw` OpenCode `unsloth start opencode` ### [](https://unsloth.ai/docs/fr#pourquoi-unsloth) 🦥 Pourquoi Unsloth ? * Nous collaborons directement avec les équipes derrière [gpt-oss](https://docs.unsloth.ai/new/gpt-oss-how-to-run-and-fine-tune#unsloth-fixes-for-gpt-oss) , [Qwen3](https://www.reddit.com/r/LocalLLaMA/comments/1kaodxu/qwen3_unsloth_dynamic_ggufs_128k_context_bug_fixes/) , [Llama 4](https://github.com/ggml-org/llama.cpp/pull/12889) , [Mistral](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/discussions/18) , [Gemma 1-3](https://news.ycombinator.com/item?id=39671146) et [Phi-4](https://unsloth.ai/blog/phi4) , où nous avons **corrigé des bugs critiques** qui ont grandement amélioré la précision du modèle. Andrej Karpathy, par exemple, a [salué notre travail](https://x.com/karpathy/status/1765473722985771335) . * Unsloth simplifie l’entraînement local, l’inférence, les données et le déploiement * Unsloth prend en charge l’inférence et l’entraînement de plus de 500 modèles : [vision](https://unsloth.ai/docs/fr/notions-de-base/vision-fine-tuning) , [TTS](https://unsloth.ai/docs/fr/notions-de-base/text-to-speech-tts-fine-tuning) , [embeddings](https://unsloth.ai/docs/fr/notions-de-base/embedding-finetuning) , [RL](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) ### [](https://unsloth.ai/docs/fr#fonctionnalites) ⭐ Fonctionnalités Unsloth vous permet d’exécuter et d’entraîner des modèles pour le texte, [audio](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning) , [embeddings](https://unsloth.ai/docs/new/embedding-finetuning) , [vision](https://unsloth.ai/docs/basics/vision-fine-tuning) et plus encore. Unsloth offre de nombreuses fonctionnalités clés pour l’inférence et l’entraînement : #### [](https://unsloth.ai/docs/fr#inference) Inférence * [Appel d’outils auto-réparateur](https://unsloth.ai/docs/fr/nouveau/studio/chat#auto-healing-tool-calling) / recherche web et utilisation [Unsloth comme API](https://unsloth.ai/docs/fr/notions-de-base/api) . * Connectez vos modèles locaux à n’importe quel agent : [Claude Code](https://unsloth.ai/docs/fr/notions-de-base/claude-code) , [Codex](https://unsloth.ai/docs/fr/notions-de-base/codex) , [Hermes](https://unsloth.ai/docs/fr/integrations/hermes-agent) et plus encore. * Recherchez + téléchargez + exécutez n’importe quel modèle, comme des GGUF, des adaptateurs LoRA, des safetensors. * [Paramètre d’inférence automatique](https://unsloth.ai/docs/fr/nouveau/studio/chat#auto-parameter-tuning) ajustement et modification des modèles de conversation. * [Exportez ou enregistrez](https://unsloth.ai/docs/fr/nouveau/studio/export) votre modèle en GGUF, en safetensor 16 bits, etc. * [Comparez les sorties](https://unsloth.ai/docs/fr/nouveau/studio/chat#model-arena) avec deux modèles différents côte à côte. #### [](https://unsloth.ai/docs/fr#entrainement) Entraînement * Entraînez et [RL](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) plus de 500 modèles ~2x plus vite avec ~70 % de VRAM en moins (sans perte de précision) * Prend en charge le fine-tuning complet, le préentraînement, et l’entraînement en 4 bits, 16 bits et FP8. * [Création automatique de jeux de données](https://unsloth.ai/docs/fr/nouveau/studio/data-recipe) à partir de fichiers PDF, CSV, DOCX. Modifiez les données dans un workflow visuel à base de nœuds. * Observabilité : surveillez l’entraînement en direct, suivez la perte, l’utilisation du GPU, personnalisez les graphiques * Le plus efficace [**apprentissage par renforcement**](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) bibliothèque, utilisant 80 % de VRAM en moins pour GRPO, [FP8](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/fp8-reinforcement-learning) etc. * [Multi-GPU](https://unsloth.ai/docs/fr/notions-de-base/multi-gpu-training-with-unsloth) fonctionne, mais une bien meilleure version arrive ! ### [](https://unsloth.ai/docs/fr#quest-ce-que-le-fine-tuning-et-le-rl-pourquoi) Qu’est-ce que le fine-tuning et le RL ? Pourquoi ? [**Fine-tuning** un LLM](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) personnalise son comportement, renforce ses connaissances dans un domaine et optimise ses performances pour des tâches spécifiques. En fine-tunant un modèle préentraîné (par ex. Llama-3.1-8B) sur un jeu de données, vous pouvez : * **Mettre à jour les connaissances** : introduire de nouvelles informations propres au domaine. * **Personnaliser le comportement** : ajuster le ton, la personnalité ou le style de réponse du modèle. * **Optimiser pour des tâches** : améliorer la précision et la pertinence pour des cas d’usage spécifiques. [**L’apprentissage par renforcement (RL)**](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) est le domaine où un « agent » apprend à prendre des décisions en interagissant avec un environnement et en recevant **des retours** sous forme de **récompenses** ou **pénalités**. * **Action :** Ce que le modèle génère (par ex. une phrase). * **Récompense :** Un signal indiquant si l’action du modèle était bonne ou mauvaise (par ex. la réponse suivait-elle les instructions ? était-elle utile ?). * **Environnement :** Le scénario ou la tâche sur laquelle le modèle travaille (par ex. répondre à la question d’un utilisateur). **Exemples de cas d’usage du fine-tuning ou du RL**: * Permet aux LLM de prédire si un titre a un impact positif ou négatif sur une entreprise. * Peut utiliser les interactions historiques des clients pour des réponses plus précises et personnalisées. * Fine-tunez un LLM sur des textes juridiques pour l’analyse de contrats, la recherche jurisprudentielle et la conformité. Vous pouvez considérer un modèle fine-tuné comme un agent spécialisé conçu pour accomplir des tâches spécifiques plus efficacement. **Le fine-tuning peut reproduire toutes les capacités du RAG** mais pas l’inverse. [Mises à jour d'Unsloth](https://unsloth.ai/docs/fr/nouveau/changelog) [🖥️Inférence et déploiement](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment) [Unsloth Start](https://unsloth.ai/docs/fr/integrations/unsloth-start) [🦥Dynamic 2.0 GGUFs](https://unsloth.ai/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-134302f2507d4313b9575917c9a43b0a0028856c%252Flarge%2520sloth%2520wave.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=9805abfe&sv=2) [SuivantModels](https://unsloth.ai/docs/fr/commencer/unsloth-model-catalog) Mis à jour il y a 2 jours Ce contenu vous a-t-il été utile ? * [⚡ Démarrage rapide](https://unsloth.ai/docs/fr#demarrage-rapide) * [👾 Unsloth Start](https://unsloth.ai/docs/fr#unsloth-start) * [🦥 Pourquoi Unsloth ?](https://unsloth.ai/docs/fr#pourquoi-unsloth) * [⭐ Fonctionnalités](https://unsloth.ai/docs/fr#fonctionnalites) * [Qu’est-ce que le fine-tuning et le RL ? Pourquoi ?](https://unsloth.ai/docs/fr#quest-ce-que-le-fine-tuning-et-le-rl-pourquoi) Ce contenu vous a-t-il été utile ? Copier unsloth start claude --- # Fine-tuning pour débutants | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners.md) . Si vous êtes débutant, voici peut-être les premières questions que vous vous poserez avant votre premier fine-tuning. Vous pouvez aussi toujours demander à notre communauté en rejoignant notre [page Reddit](https://www.reddit.com/r/unsloth/) . [](https://unsloth.ai/docs/fr/modeles/tutorials) [Complete LLM Directory](https://unsloth.ai/docs/fr/modeles/tutorials) Découvrez tous les modèles que vous pouvez exécuter / entraîner avec Unsloth. Comment exécuter des GGUF ou entraîner des LLM ? [](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) 🧬[Fine-tuning Guide](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) Étape par étape sur la façon de faire du fine-tuning ! Apprenez les bases fondamentales de l'entraînement. [](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/what-model-should-i-use) ❓[What Model Should I Use?](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/what-model-should-i-use) Modèle Instruct ou modèle de base ? Quelle taille devrait avoir mon jeu de données ? [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me) 🤔[FAQ + Le fine-tuning est-il fait pour moi ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me) Que peut m'apporter le fine-tuning ? RAG vs. fine-tuning ? [](https://unsloth.ai/docs/fr/commencer/install) 📥[Installation](https://unsloth.ai/docs/fr/commencer/install) Comment installer Unsloth localement ? Comment mettre à jour Unsloth ? 📈[Guide des jeux de données](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/datasets-guide) Comment structurer/préparer mon jeu de données ? Comment collecter des données ? [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) 🛠️[Exigences Unsloth](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) Unsloth fonctionne-t-il sur mon GPU ? De quelle quantité de VRAM aurai-je besoin ? [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment) 🖥️[Inférence et déploiement](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment) Comment enregistrer mon modèle localement ? Comment exécuter mon modèle via Ollama ou vLLM ? 🧠[Hyperparameters Guide](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide) Que se passe-t-il lorsque je modifie un paramètre ? Quels paramètres dois-je modifier ? ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-559e7890f607e34fd6004517296e65e942c93b41%252FLarge%2520sloth%2520Question%2520mark.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=29182802&sv=2) [PrécédentModels](https://unsloth.ai/docs/fr/commencer/unsloth-model-catalog) [SuivantExigences Unsloth](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) Mis à jour il y a 2 mois Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? --- # Hackathon de Reinforcement Learning IA AMD avec Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/amd/amd-hackathon.md) . Vous pouvez consulter le dépôt GitHub d'Unsloth ici : [https://github.com/unslothai/unsloth](https://github.com/unslothai/unsloth) Voici le lien vers nos notebooks de fine-tuning AMD : [![Logo](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2Fgithub.com%2Ffluidicon.png&width=20&dpr=3&quality=100&sign=dbc3edae&sv=2)notebooks/nb/gpt\_oss\_(20B)\_Reinforcement\_Learning\_2048\_Game\_BF16.ipynb at main · unslothai/notebooksGitHub](https://github.com/unslothai/notebooks/blob/main/nb/gpt_oss_(20B)_Reinforcement_Learning_2048_Game_BF16.ipynb) [https://github.com/unslothai/notebooks/blob/main/nb/gpt\_oss\_(20B)\_Reinforcement\_Learning\_2048\_Game\_BF16.ipynb](https://github.com/unslothai/notebooks/blob/main/nb/gpt_oss_(20B)_Reinforcement_Learning_2048_Game_BF16.ipynb) Si vous souhaitez mettre à jour Unsloth / Unsloth Zoo : Pour bitsandbytes : Si vous voyez : N'utilisez PAS UV\_SKIP\_WHEEL\_FILENAME\_CHECK, utilisez plutôt UNIQUEMENT `pip install "unsloth[amd] @ git+https://github.com/unslothai/unsloth"` (PAS uv) car uv détruit bitsandbytes. Peut-être ajouter une vérification dans les PR si possible pour détecter cela. Pour les instructions d'installation AMD, vous pouvez consulter notre guide ici : [AMD](https://unsloth.ai/docs/fr/commencer/install/amd) Mis à jour il y a 5 mois Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? Copier wget 'https://raw.githubusercontent.com/unslothai/notebooks/refs/heads/main/nb/gpt_oss_(20B)_Reinforcement_Learning_2048_Game_BF16.ipynb' Copier uv pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm7.0 --upgrade --force-reinstall pip uninstall unsloth unsloth_zoo -y && \ pip install git+https://github.com/unslothai/unsloth-zoo git+https://github.com/unslothai/unsloth --no-deps --force-reinstall --no-cache-dir Copier pip install "unsloth[amd] @ git+https://github.com/unslothai/unsloth" Copier error: Failed to install: bitsandbytes-1.33.7rc0-py3-none-manylinux_2_24_x86_64.whl (bitsandbytes==1.33.7rc0 (from https://github.com/bitsandbytes-foundation/bitsandbytes/releases/download/continuous-release_main/bitsandbytes-1.33.7.preview-py3-none-manylinux_2_24_x86_64.whl)) Caused by: Wheel version does not match filename (0.49.2.dev0 != 1.33.7rc0), which indicates a malformed wheel. If this is intentional, set UV_SKIP_WHEEL_FILENAME_CHECK=1. --- # Installation via Conda | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/conda-install.md) . N'utilisez Conda que si vous l'avez. Sinon, utilisez [Pip](https://unsloth.ai/docs/fr/commencer/install/pip-install) . Sélectionnez soit `pytorch-cuda=11.8,12.1` pour CUDA 11.8 ou CUDA 12.1. Nous prenons en charge `python=3.10,3.11,3.12`. Copier conda create --name unsloth_env python=3.11 -y conda activate unsloth_env pip install unsloth Si vous souhaitez installer Conda dans un environnement Linux, [lisez ici](https://docs.anaconda.com/miniconda/) , ou exécutez ce qui suit : Copier mkdir -p ~/miniconda3 wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -O ~/miniconda3/miniconda.sh bash ~/miniconda3/miniconda.sh -b -u -p ~/miniconda3 rm -rf ~/miniconda3/miniconda.sh ~/miniconda3/bin/conda init bash ~/miniconda3/bin/conda init zsh Mis à jour il y a 5 mois Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? --- # Mettre à jour Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/updating.md) . ### [](https://unsloth.ai/docs/fr/commencer/install/updating#mettre-a-jour-unsloth-studio) **Mettre à jour Unsloth Studio** Vous pouvez utiliser les mêmes commandes d'installation pour mettre à jour #### [](https://unsloth.ai/docs/fr/commencer/install/updating#macos-linux-wsl) **MacOS, Linux, WSL :** Copier curl -fsSL https://unsloth.ai/install.sh | sh #### [](https://unsloth.ai/docs/fr/commencer/install/updating#windows-powershell) **Windows PowerShell :** Copier irm https://unsloth.ai/install.ps1 | iex ### [](https://unsloth.ai/docs/fr/commencer/install/updating#mise-a-jour-du-coeur-dunsloth) Mise à jour du cœur d’Unsloth : Copier pip install --upgrade unsloth unsloth_zoo #### [](https://unsloth.ai/docs/fr/commencer/install/updating#mise-a-jour-du-coeur-dunsloth-sans-mise-a-jour-des-dependances) Mise à jour du cœur d’Unsloth sans mise à jour des dépendances : Copier pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth_zoo #### [](https://unsloth.ai/docs/fr/commencer/install/updating#pour-utiliser-une-ancienne-version-dunsloth) Pour utiliser une ancienne version d’Unsloth : Copier pip install --force-reinstall --no-cache-dir --no-deps unsloth==2025.1.5 '2025.1.5' est l'une des anciennes versions d’Unsloth. Remplacez-la par une version spécifique répertoriée sur notre [GitHub ici](https://github.com/unslothai/unsloth/releases) . [PrécédentDocker](https://unsloth.ai/docs/fr/commencer/install/docker) [SuivantIntel](https://unsloth.ai/docs/fr/commencer/install/intel) Mis à jour il y a 1 mois Ce contenu vous a-t-il été utile ? * [Mettre à jour Unsloth Studio](https://unsloth.ai/docs/fr/commencer/install/updating#mettre-a-jour-unsloth-studio) * [Mise à jour du cœur d’Unsloth :](https://unsloth.ai/docs/fr/commencer/install/updating#mise-a-jour-du-coeur-dunsloth) Ce contenu vous a-t-il été utile ? --- # Google Colab | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/google-colab.md) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-4d1b1778f3c8bde62a40130d7b4395b8bb1ce90f%252FColab%2520Options.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=27d3fc4e&sv=2) Si vous n'avez jamais utilisé un carnet Colab, un bref aperçu du carnet lui-même : 1. **Bouton Lecture sur chaque "cellule".** Cliquez sur ceci pour exécuter le code de cette cellule. Vous ne devez sauter aucune cellule et vous devez exécuter chaque cellule dans l'ordre chronologique. Si vous rencontrez des erreurs, réexécutez simplement la cellule que vous n'avez pas exécutée. Une autre option est d'appuyer sur CTRL + ENTER si vous ne voulez pas cliquer sur le bouton de lecture. 2. **Bouton Runtime dans la barre d'outils en haut.** Vous pouvez également utiliser ce bouton et cliquer sur "Exécuter tout" pour lancer l'ensemble du notebook en une seule fois. Cela sautera toutes les étapes de personnalisation, mais c'est un bon premier essai. 3. **Bouton Se connecter / Reconnecter T4.** Le T4 est le GPU gratuit fourni par Google. Il est assez puissant ! La première cellule d'installation ressemble à ceci : N'oubliez pas de cliquer sur le bouton LIRE dans les crochets \[ \]. Nous récupérons notre package open source sur Github et installons quelques autres packages. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-135cd796b01420cc4d5ce3ca243e9065154070a5%252Fimage%2520%2813%29%2520%281%29%2520%281%29.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=96cbeae5&sv=2) ### [](https://unsloth.ai/docs/fr/commencer/install/google-colab#exemple-de-code-colab) Exemple de code Colab Exemple de code Unsloth pour ajuster finement gpt-oss-20b : Mis à jour il y a 6 mois Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? Copier from unsloth import FastLanguageModel, FastModel import torch from trl import SFTTrainer, SFTConfig #token = "hf_...", # en utiliser un si vous utilisez des modèles régulés comme meta-llama/Llama-2-7b-hf max_seq_length = 2048 # Prend en charge l'évolutivité RoPE en interne, choisissez donc n'importe quelle valeur ! # Obtenir le dataset LAION url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified_chip2.jsonl" dataset = load_dataset("json", data_files = {"train" : url}, split = "train") # Modèles pré-quantifiés 4 bits que nous prenons en charge pour un téléchargement 4× plus rapide + pas d'OOM. fourbit_models = [\ "unsloth/gpt-oss-20b-unsloth-bnb-4bit", # ou choisissez n'importe quel modèle\ \ ] # Plus de modèles sur https://huggingface.co/unsloth model, tokenizer = FastModel.from_pretrained( model_name = "unsloth/gpt-oss-20b", max_seq_length = 2048, # Choisissez n'importe quoi pour un long contexte ! load_in_4bit = True, # quantification 4 bits. False = LoRA 16 bits. load_in_8bit = False, # quantification 8 bits load_in_16bit = False, # [NOUVEAU !] LoRA 16 bits full_finetuning = False, # À utiliser pour un affinement complet. # token = "hf_...", # utilisez-en un si vous utilisez des modèles à accès restreint ) # Effectuer le patching du modèle et ajouter des poids LoRA rapides model = FastLanguageModel.get_peft_model( model, r = 16, target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",\ "gate_proj", "up_proj", "down_proj",], lora_alpha = 16, lora_dropout = 0, # Prend en charge n'importe quelle valeur, mais = 0 est optimisé bias = "none", # Prend en charge n'importe quelle valeur, mais = "none" est optimisé # [NOUVEAU] "unsloth" utilise 30% de VRAM en moins, permet des tailles de batch 2× supérieures ! use_gradient_checkpointing = "unsloth", # True or "unsloth" pour des contextes très longs random_state = 3407, max_seq_length = max_seq_length, use_rslora = False, # Nous prenons en charge le LoRA à rang stabilisé loftq_config = None, # Et LoftQ ) trainer = SFTTrainer( trainer = Trainer( model = model, tokenizer = tokenizer, args = SFTConfig( max_seq_length = max_seq_length, per_device_train_batch_size = 2, gradient_accumulation_steps = 4, warmup_steps = 10, # num_train_epochs = 1, # Définissez ceci pour 1 cycle complet d'entraînement. bf16 = is_bfloat16_supported(), output_dir = "outputs", optim = "adamw_8bit", seed = 3407, ), ) trainer.train() # Allez sur https://docs.unsloth.ai pour des astuces avancées comme # (1) Sauvegarder en GGUF / fusionner en 16 bits pour vLLM # (2) Poursuivre l'entraînement à partir d'un adaptateur LoRA sauvegardé # (3) Ajouter une boucle d'évaluation / OOMs # (4) Modèles de chat personnalisés --- # FAQ + Le fine-tuning est-il fait pour moi ? | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me.md) . [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#comprendre-laffinage-fine-tuning) Comprendre l’affinage (Fine-Tuning) ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ L’affinage d’un LLM personnalise son comportement, approfondit son expertise de domaine et optimise ses performances pour des tâches spécifiques. En affinant un modèle pré-entraîné (par ex. _Llama-3.1-8B_) avec des données spécialisées, vous pouvez : * **Mettre à jour les connaissances** – Introduire de nouvelles informations spécifiques au domaine que le modèle de base n’incluait pas à l’origine. * **Personnaliser le comportement** – Ajuster le ton, la personnalité ou le style de réponse du modèle pour répondre à des besoins spécifiques ou à la voix d’une marque. * **Optimiser pour des tâches** – Améliorer la précision et la pertinence sur des tâches ou requêtes particulières requises par votre cas d’usage. Pensez à l’affinage comme à la création d’un expert spécialisé à partir d’un modèle généraliste. Certains débattent de l’utilisation de la génération augmentée par récupération (RAG) au lieu de l’affinage, mais l’affinage peut intégrer des connaissances et des comportements directement dans le modèle de manières que la RAG ne peut pas. En pratique, combiner les deux approches donne les meilleurs résultats - conduisant à une plus grande précision, une meilleure utilisabilité et moins d’hallucinations. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#applications-reelles-de-laffinage) Applications réelles de l’affinage L’affinage peut être appliqué dans divers domaines et répondre à différents besoins. Voici quelques exemples pratiques de la différence qu’il peut faire : * **Analyse de sentiment pour la finance** – Entraîner un LLM à déterminer si un titre d’actualité affecte une entreprise positivement ou négativement, en adaptant sa compréhension au contexte financier. * **Chatbots de support client** – Affiner sur les interactions clients passées pour fournir des réponses plus précises et personnalisées dans le style et la terminologie de l’entreprise. * **Assistance pour documents juridiques** – Affiner sur des textes juridiques (contrats, jurisprudence, réglementations) pour des tâches comme l’analyse de contrats, la recherche de jurisprudence ou le support conformité, en veillant à ce que le modèle utilise un langage juridique précis. [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#les-avantages-de-laffinage) Les avantages de l’affinage ---------------------------------------------------------------------------------------------------------------------------------------------------------------- L’affinage offre plusieurs avantages notables au-delà de ce qu’un modèle de base ou un système purement basé sur la récupération peut fournir : #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#affinage-vs-rag-quelle-est-la-difference) Affinage vs RAG : Quelle est la différence ? L’affinage peut faire à peu près tout ce que la RAG peut faire - mais pas l’inverse. Pendant l’entraînement, l’affinage intègre les connaissances externes directement dans le modèle. Cela permet au modèle de traiter des requêtes de niche, de résumer des documents et de maintenir le contexte sans dépendre d’un système de récupération externe. Cela ne veut pas dire que la RAG manque d’avantages, car elle excelle pour accéder à des informations à jour depuis des bases externes. Il est en fait possible de récupérer des données fraîches avec l’affinage également, cependant il est préférable de combiner la RAG avec l’affinage pour plus d’efficacité. #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#maitrise-specifique-a-la-tache) Maîtrise spécifique à la tâche L’affinage intègre profondément les connaissances de domaine dans le modèle. Cela le rend très efficace pour gérer des requêtes structurées, répétitives ou nuancées, des scénarios où les systèmes uniquement basés sur la RAG ont souvent du mal. En d’autres termes, un modèle affiné devient un spécialiste des tâches ou du contenu sur lesquels il a été entraîné. #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#independance-de-la-recuperation) Indépendance de la récupération Un modèle affiné n’a pas de dépendance aux sources de données externes au moment de l’inférence. Il reste fiable même si un système de récupération connecté échoue ou est incomplet, car toutes les informations nécessaires sont déjà contenues dans les propres paramètres du modèle. Cette autosuffisance signifie moins de points de défaillance en production. #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#reponses-plus-rapides) Réponses plus rapides Les modèles affinés n’ont pas besoin d’appeler une base de connaissances externe pendant la génération. Sauter l’étape de récupération leur permet de produire des réponses beaucoup plus rapidement. Cette rapidité rend les modèles affinés idéaux pour des applications sensibles au temps où chaque seconde compte. #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#comportement-et-ton-personnalises) Comportement et ton personnalisés L’affinage permet un contrôle précis sur la façon dont le modèle communique. Cela assure que les réponses du modèle restent cohérentes avec la voix d’une marque, respectent les exigences réglementaires ou correspondent à des préférences de ton spécifiques. Vous obtenez un modèle qui non seulement sait _quoi_ dire, mais _comment_ le dire dans le style souhaité. #### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#performance-fiable) Performance fiable Même dans une configuration hybride qui utilise à la fois l’affinage et la RAG, le modèle affiné fournit une solution de repli fiable. Si le composant de récupération n’arrive pas à trouver la bonne information ou retourne des données incorrectes, les connaissances intégrées du modèle peuvent toujours générer une réponse utile. Cela garantit des performances plus cohérentes et robustes pour votre système. [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#idees-recues-courantes) Idées reçues courantes ------------------------------------------------------------------------------------------------------------------------------------------------------- Malgré les avantages de l’affinage, quelques mythes persistent. Abordons deux des idées reçues les plus courantes concernant l’affinage : ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#laffinage-ajoute-t-il-de-nouvelles-connaissances-a-un-modele) L’affinage ajoute-t-il de nouvelles connaissances à un modèle ? **Oui - il le peut absolument.** Un mythe courant suggère que l’affinage n’introduit pas de nouvelles connaissances, mais en réalité il le fait. Si votre jeu de données d’affinage contient des informations nouvelles et spécifiques au domaine, le modèle apprendra ce contenu pendant l’entraînement et l’incorporera dans ses réponses. En effet, l’affinage _peut et fait_ apprendre au modèle de nouveaux faits et schémas depuis zéro. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#la-rag-est-elle-toujours-meilleure-que-laffinage) La RAG est-elle toujours meilleure que l’affinage ? **Pas nécessairement.** Beaucoup supposent que la RAG surpassera systématiquement un modèle affiné, mais ce n’est pas le cas lorsque l’affinage est bien réalisé. En fait, un modèle bien affiné égalise souvent voire dépasse les systèmes basés sur la RAG sur des tâches spécialisées. Les affirmations selon lesquelles « la RAG est toujours meilleure » proviennent généralement de tentatives d’affinage mal configurées - par exemple, l’utilisation de [paramètres LoRA](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide) incorrects ou d’un entraînement insuffisant. Unsloth se charge de ces complexités en sélectionnant automatiquement les meilleures configurations de paramètres pour vous. Tout ce dont vous avez besoin est un jeu de données de bonne qualité, et vous obtiendrez un modèle affiné qui donne le meilleur de ses capacités. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#laffinage-est-il-cher) L’affinage est-il cher ? **Pas du tout !** Alors que l’affinage complet ou le pré-entraînement peut être coûteux, ceux-ci ne sont pas nécessaires (le pré-entraînement n’est particulièrement pas nécessaire). Dans la plupart des cas, l’affinage LoRA ou QLoRA peut être effectué à coût minimal. En fait, avec les [notebooks gratuits](https://docs.unsloth.ai/get-started/unsloth-notebooks) d’Unsloth pour Colab ou Kaggle, vous pouvez affiner des modèles sans dépenser un centime. Mieux encore, vous pouvez même affiner localement sur votre propre appareil. [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#faq) FAQ : ------------------------------------------------------------------------------------------------------------------- ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#pourquoi-vous-devriez-combiner-rag-et-affinage) Pourquoi vous devriez combiner RAG et affinage Au lieu de choisir entre la RAG et l’affinage, envisagez d’utiliser **les deux** ensemble pour de meilleurs résultats. Combiner un système de récupération avec un modèle affiné fait ressortir les points forts de chaque approche. Voici pourquoi : * **Expertise spécifique à la tâche** – L’affinage excelle dans les tâches ou formats spécialisés (faisant du modèle un expert dans un domaine précis), tandis que la RAG maintient le modèle à jour avec les dernières connaissances externes. * **Meilleure adaptabilité** – Un modèle affiné peut toujours fournir des réponses utiles même si le composant de récupération échoue ou renvoie des informations incomplètes. Pendant ce temps, la RAG assure que le système reste actuel sans exiger que vous réentraîniez le modèle pour chaque nouvelle donnée. * **Efficacité** – L’affinage fournit une solide base de connaissances intégrée au modèle, et la RAG gère les détails dynamiques ou rapidement changeants sans nécessiter un réentraînement exhaustif depuis zéro. Cet équilibre offre un flux de travail efficace et réduit les coûts de calcul globaux. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#lora-vs-qlora-lequel-utiliser) LoRA vs QLoRA : Lequel utiliser ? Lorsqu’il s’agit de mettre en œuvre l’affinage, deux techniques populaires peuvent réduire considérablement les besoins en calcul et en mémoire : **LoRA** et **QLoRA**. Voici une brève comparaison de chacune : * **LoRA (Low-Rank Adaptation)** – Affine uniquement un petit ensemble de matrices de poids supplémentaires “adaptateur” (en précision 16 bits), tout en laissant la plupart du modèle original inchangé. Cela réduit considérablement le nombre de paramètres devant être mis à jour pendant l’entraînement. * **QLoRA (Quantized LoRA)** – Combine LoRA avec une quantification 4 bits des poids du modèle, permettant un affinage efficace de très grands modèles sur du matériel minimal. En utilisant la précision 4 bits là où c’est possible, elle réduit fortement l’utilisation de la mémoire et la charge de calcul. Nous recommandons de commencer par **QLoRA**, car c’est l’une des méthodes les plus efficaces et accessibles disponibles. Grâce aux [quants dynamiques 4 bits](https://unsloth.ai/blog/dynamic-4bit) d’Unsloth, la perte d’exactitude comparée à un affinage LoRA standard en 16 bits est désormais négligeable. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#lexperimentation-est-essentielle) L'expérimentation est essentielle Il n’y a pas d’approche « meilleure » unique pour l’affinage - seulement des bonnes pratiques selon les scénarios. Il est important d’expérimenter différentes méthodes et configurations pour trouver ce qui fonctionne le mieux pour votre jeu de données et votre cas d’usage. Un excellent point de départ est **QLoRA (4 bits)**, qui offre une manière très rentable et peu gourmande en ressources d’affiner des modèles sans exigences computationnelles lourdes. [🧠Hyperparameters Guide](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide) [PrécédentExigences Unsloth](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) [SuivantNotebooks Unsloth](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks) Mis à jour il y a 6 mois Ce contenu vous a-t-il été utile ? * [Comprendre l’affinage (Fine-Tuning)](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#comprendre-laffinage-fine-tuning) * [Applications réelles de l’affinage](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#applications-reelles-de-laffinage) * [Les avantages de l’affinage](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#les-avantages-de-laffinage) * [Idées reçues courantes](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#idees-recues-courantes) * [L’affinage ajoute-t-il de nouvelles connaissances à un modèle ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#laffinage-ajoute-t-il-de-nouvelles-connaissances-a-un-modele) * [La RAG est-elle toujours meilleure que l’affinage ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#la-rag-est-elle-toujours-meilleure-que-laffinage) * [L’affinage est-il cher ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#laffinage-est-il-cher) * [FAQ :](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#faq) * [Pourquoi vous devriez combiner RAG et affinage](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#pourquoi-vous-devriez-combiner-rag-et-affinage) * [LoRA vs QLoRA : Lequel utiliser ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#lora-vs-qlora-lequel-utiliser) * [L'expérimentation est essentielle](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me#lexperimentation-est-essentielle) Ce contenu vous a-t-il été utile ? --- # Installation d'Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install.md) . Unsloth peut être utilisé de deux façons : via [Unsloth Studio](https://unsloth.ai/docs/fr/nouveau/studio/install) , l'interface web, ou via Unsloth Core, la version originale basée sur le code. Consultez nos [exigences système](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements) Unsloth Studio fonctionne sur MacOS, Linux, Windows, NVIDIA, et plus encore. Utilisez les mêmes commandes d'installation pour mettre à jour. **MacOS, Linux, WSL :** Copier curl -fsSL https://unsloth.ai/install.sh | sh **Windows PowerShell :** Copier irm https://unsloth.ai/install.ps1 | iex **Lancez Unsloth Studio :** Copier unsloth studio -H 0.0.0.0 -p 8888 [MacOS](https://unsloth.ai/docs/fr/commencer/install/mac) [](https://unsloth.ai/docs/fr/commencer/install/pip-install) [uv, pip install & venv](https://unsloth.ai/docs/fr/commencer/install/pip-install) [Windows](https://unsloth.ai/docs/fr/commencer/install/windows-installation) [AMD](https://unsloth.ai/docs/fr/commencer/install/amd) [](https://unsloth.ai/docs/fr/commencer/install/updating) [Updating](https://unsloth.ai/docs/fr/commencer/install/updating) [Docker](https://unsloth.ai/docs/fr/commencer/install/docker) [Intel](https://unsloth.ai/docs/fr/commencer/install/intel) [](https://unsloth.ai/docs/fr/commencer/install/conda-install) [Conda](https://unsloth.ai/docs/fr/commencer/install/conda-install) [VS Code](https://unsloth.ai/docs/fr/commencer/install/vs-code) [](https://unsloth.ai/docs/fr/commencer/install/google-colab) [Google Colab](https://unsloth.ai/docs/fr/commencer/install/google-colab) [PrécédentNotebooks Unsloth](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks) [Suivantuv, pip install & venv](https://unsloth.ai/docs/fr/commencer/install/pip-install) Mis à jour il y a 2 jours Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? --- # Installer Unsloth sur MacOS | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/mac.md) . Pour installer Unsloth localement sur votre appareil Apple MacOS local, suivez les étapes ci-dessous : ### [](https://unsloth.ai/docs/fr/commencer/install/mac#installer-unsloth) Installer Unsloth Copier curl -fsSL https://unsloth.ai/install.sh | sh Utilisez la même commande pour **mettre à jour**. ### [](https://unsloth.ai/docs/fr/commencer/install/mac#lancer) Lancer Chaque fois que vous souhaitez relancer Unsloth : Copier unsloth studio -H 0.0.0.0 -p 8888 Pour des instructions d'installation détaillées et les prérequis d'Unsloth Studio, [consultez notre guide](https://unsloth.ai/docs/fr/nouveau/studio/install) . ### [](https://unsloth.ai/docs/fr/commencer/install/mac#desinstallation) Désinstallation La manière recommandée de supprimer complètement Unsloth Studio est d'utiliser le script de désinstallation correspondant à votre OS. Il arrête tous les serveurs en cours d'exécution, supprime l'application, la commande CLI, les données du lanceur, les raccourcis et les entrées spécifiques à la plateforme (macOS `.app` bundle + Launch Services ; menu Démarrer Windows + registre + PATH) : Copier curl -fsSL https://raw.githubusercontent.com/unslothai/unsloth/main/scripts/uninstall.sh | sh #### [](https://unsloth.ai/docs/fr/commencer/install/mac#desinstallation-manuelle) Désinstallation manuelle Si vous préférez supprimer uniquement certaines parties : **1\. Supprimer uniquement l'application** (conserve l'historique, les discussions, les points de contrôle et les exportations intacts) : * `rm -rf ~/.unsloth/studio/unsloth_studio` **2\. Supprimer complètement Unsloth** (conserve intacts les autres outils Unsloth) : * `rm -rf ~/.unsloth/studio` **3\. Supprimer tout ce qui est lié à Unsloth :** * `rm -rf ~/.unsloth` Remarque : l'étape 3 supprime tout l'historique, les discussions, les points de contrôle du modèle et les exportations. Cela ne peut pas être annulé. **4\. Supprimer les raccourcis et les liens symboliques :** **5\. Supprimer la commande CLI :** * `rm -f ~/.local/bin/unsloth` Remarque : les étapes 1 à 5 ne touchent pas aux fichiers de modèle HF que vous avez téléchargés. Voir la section Suppression des fichiers de modèle HF mis en cache ci-dessous si vous souhaitez récupérer cet espace. ### [](https://unsloth.ai/docs/fr/commencer/install/mac#suppression-des-fichiers-de-modele) **Suppression des fichiers de modèle** Vous pouvez supprimer les anciens fichiers de modèle soit via l'icône de corbeille dans la recherche de modèles, soit en supprimant le dossier de modèle mis en cache correspondant dans le répertoire de cache Hugging Face. L'emplacement de cache par défaut est : Si `HF_HUB_CACHE` ou `HF_HOME` est défini, utilisez plutôt cet emplacement. Vous pouvez vérifier avec : Pour supprimer un modèle spécifique, supprimez son dossier (par ex. `models--unsloth--Llama-3.1-8B-bnb-4bit`) du répertoire de cache. Pour vider tous les modèles mis en cache : [Précédentuv, pip install & venv](https://unsloth.ai/docs/fr/commencer/install/pip-install) [SuivantWindows](https://unsloth.ai/docs/fr/commencer/install/windows-installation) Mis à jour il y a 4 jours Ce contenu vous a-t-il été utile ? * [Installer Unsloth](https://unsloth.ai/docs/fr/commencer/install/mac#installer-unsloth) * [Lancer](https://unsloth.ai/docs/fr/commencer/install/mac#lancer) * [Désinstallation](https://unsloth.ai/docs/fr/commencer/install/mac#desinstallation) * [Suppression des fichiers de modèle](https://unsloth.ai/docs/fr/commencer/install/mac#suppression-des-fichiers-de-modele) Ce contenu vous a-t-il été utile ? Copier rm -rf ~/Applications/Unsloth\ Studio.app ~/Desktop/Unsloth\ Studio Copier ~/.cache/huggingface/hub/ Copier echo ${HF_HUB_CACHE:-${HF_HOME:-${XDG_CACHE_HOME:-$HOME/.cache}/huggingface}/hub} Copier rm -rf ~/.cache/huggingface/hub/ --- # Exigences Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements.md) . Unsloth peut être utilisé de deux façons : via [Unsloth Studio](https://unsloth.ai/docs/fr/nouveau/studio/install) , l’interface web, ou via [Unsloth Core](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#unsloth-core-requirements) , la version originale basée sur le code. Chacun a des exigences différentes. [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-dunsloth-studio) **Exigences d’Unsloth Studio** ----------------------------------------------------------------------------------------------------------------------------------------------------- * **Mac :** L’entraînement, MLX et l’inférence GGUF sont TOUS pris en charge. * **CPU : Unsloth fonctionne toujours sans GPU**, pour Chat + les recettes de données. * **Entraînement :** Fonctionne sur **NVIDIA**, **AMD**, **Intel** GPU et **Mac** appareils ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#windowss) Windows**s** Unsloth Studio fonctionne directement sur Windows sans WSL. Pour entraîner des modèles, assurez-vous que votre système satisfait à ces exigences : **Prérequis** * Windows 10 ou Windows 11 (64 bits) * GPU NVIDIA avec les pilotes installés * **App Installer** (inclut `winget`): [ici](https://learn.microsoft.com/en-us/windows/msix/app-installer/install-update-app-installer) * **Git**: `winget install --id Git.Git -e --source winget` * **Python** : version 3.11 jusqu’à, mais sans inclure, 3.14 * Travaillez dans un environnement Python tel que **uv**, **venv**, ou **conda/mamba** ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#macos) macOS Unsloth Studio fonctionne sur les appareils macOS avec une prise en charge complète de l’entraînement, de MLX et de l’inférence GGUF. * macOS 12 Monterey ou plus récent (Intel ou Apple Silicon) * Installez Homebrew : [ici](https://brew.sh/) * Git : `brew install git` * cmake : `brew install cmake` * openssl : `brew install openssl` * Python : version 3.11 jusqu’à, mais sans inclure, 3.14 * Travaillez dans un environnement Python tel que **uv**, **venv**, ou **conda/mamba** ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#linux-et-wsl) Linux et WSL * Ubuntu 20.04+ ou distribution similaire (64 bits) * GPU NVIDIA avec les pilotes installés * Boîte à outils CUDA (12.4+ recommandé, 12.8+ pour Blackwell) * Git : `sudo apt install git` * Python : version 3.11 jusqu’à, mais sans inclure, 3.14 * Travaillez dans un environnement Python tel que **uv**, **venv**, ou **conda/mamba** ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#cpu-uniquement) CPU uniquement Unsloth Studio prend en charge les appareils CPU pour [Chat](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#run-models-locally) pour les modèles GGUF et [Recettes de données](https://unsloth.ai/docs/fr/nouveau/studio/data-recipe) ([Export](https://unsloth.ai/docs/fr/nouveau/studio/export) (bientôt disponible) * Identiques à ceux mentionnés ci-dessus pour Linux (sauf les pilotes GPU NVIDIA) et macOS. ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#entrainement) **Entraînement** L’entraînement d’Unsloth Studio fonctionne actuellement sur NVIDIA, [AMD](https://unsloth.ai/docs/fr/commencer/install/amd) , MLX, et les appareils Intel. Vous pouvez toujours utiliser [Unsloth Core original](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#unsloth-requirements) pour entraîner sur des appareils AMD et Intel. **Python 3.11–3.13** est requis. Exigence Linux / WSL Windows **Git** Généralement préinstallé Installé par le script d’installation (`winget`) **CMake** Préinstallé ou `sudo apt install cmake` Installé par le script d’installation (`winget`) **Compilateur C++** `build-essential` Outils de build Visual Studio 2022 **CUDA Toolkit** Facultatif ; `nvcc` détecté automatiquement Installé par le script d’installation (adapté au pilote) [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-dunsloth-core) Exigences d’Unsloth Core --------------------------------------------------------------------------------------------------------------------------------------------- * **Système d’exploitation** : fonctionne sur Linux et [Windows](https://docs.unsloth.ai/get-started/install-and-update/windows-installation) * Prend en charge les GPU NVIDIA depuis 2018+, y compris [Blackwell RTX 50](https://unsloth.ai/docs/fr/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth) et [DGX Spark](https://unsloth.ai/docs/fr/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth) * Capacité CUDA minimale 7.0 (V100, T4, Titan V, RTX 20 et 50, A100, H100, L40, etc.) [Vérifiez votre GPU !](https://developer.nvidia.com/cuda-gpus) Les GTX 1070 et 1080 fonctionnent, mais sont lentes. * L’image Docker officielle [Docker d’Unsloth](https://hub.docker.com/r/unsloth/unsloth) `unsloth/unsloth` est disponible sur Docker Hub * [Docker](https://unsloth.ai/docs/fr/commencer/install/docker) * Unsloth fonctionne sur [AMD](https://unsloth.ai/docs/fr/commencer/install/amd) et [Intel](https://unsloth.ai/docs/fr/commencer/install/intel) GPU (suivez nos [guides spécifiques](https://unsloth.ai/docs/fr/commencer/install) ). Apple/Silicon/MLX est en préparation * Votre appareil devrait disposer de `xformers`, `torch`, `BitsandBytes` et `triton` prise en charge. * Si vous avez différentes versions de torch, transformers, etc., `pip install unsloth` installera automatiquement toutes les dernières versions de ces bibliothèques, donc vous n’avez pas à vous soucier de la compatibilité des versions. Python 3.13 est pris en charge ! ### [](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-vram-pour-le-fine-tuning) Exigences VRAM pour le fine-tuning : De combien de mémoire GPU ai-je besoin pour le fine-tuning d’un LLM avec Unsloth ? Un problème courant lorsque vous êtes en OOM ou à court de mémoire vient du fait que vous avez défini une taille de batch trop élevée. Réglez-la sur 1, 2 ou 3 pour utiliser moins de VRAM. **Pour les benchmarks de longueur de contexte, voir** [**ici**](https://unsloth.ai/docs/fr/notions-de-base/unsloth-benchmarks#context-length-benchmarks) **.** Consultez ce tableau pour les exigences en VRAM, classées par paramètres du modèle et méthode de fine-tuning. QLoRA utilise du 4 bits, LoRA utilise du 16 bits. Gardez à l’esprit que, parfois, davantage de VRAM est nécessaire selon le modèle, donc ces valeurs sont le minimum absolu : Paramètres du modèle VRAM QLoRA (4 bits) VRAM LoRA (16 bits) 3B 3,5 Go 8 Go 7B 5 Go 19 Go 8B 6 Go 22 Go 9B 6,5 Go 24 Go 11B 7,5 Go 29 Go 14B 8,5 Go 33 Go 27B 22 Go 64 Go 32B 26 Go 76 Go 40B 30 Go 96 Go 70B 41 Go 164 Go 81B 48 Go 192 Go 90B 53 Go 212 Go 405B 237 Go 950 Go [PrécédentBeginner?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners) [SuivantFAQ + Le fine-tuning est-il fait pour moi ?](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me) Mis à jour il y a 4 jours Ce contenu vous a-t-il été utile ? * [Exigences d’Unsloth Studio](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-dunsloth-studio) * [Windowss](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#windowss) * [macOS](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#macos) * [Linux et WSL](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#linux-et-wsl) * [CPU uniquement](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#cpu-uniquement) * [Entraînement](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#entrainement) * [Exigences d’Unsloth Core](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-dunsloth-core) * [Exigences VRAM pour le fine-tuning :](https://unsloth.ai/docs/fr/commencer/fine-tuning-for-beginners/unsloth-requirements#exigences-vram-pour-le-fine-tuning) Ce contenu vous a-t-il été utile ? --- # Comment fine-tuner des LLM dans VS Code avec Unsloth et les GPU Colab | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/vs-code.md) . Vous pouvez désormais affiner des LLM directement depuis Visual Studio Code (VS Code), localement ou en utilisant l'extension Google Colab. Dans ce guide, vous apprendrez à utiliser l'entraînement open source [dépôt : Unsloth](https://github.com/unslothai/unsloth) , pour connecter n'importe quel [carnet de fine-tuning](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks) dans VS Code à un runtime Colab, afin que vous puissiez entraîner sur votre GPU local ou sur le GPU gratuit de Colab. Vous pouvez aussi regarder notre tutoriel vidéo [ici](https://unsloth.ai/docs/fr/commencer/install/vs-code#video-tutorial) . 1 ### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#tutoriel-vs-code-et-colab) Tutoriel VS Code et Colab : Pour commencer, nous aurons besoin de : * Installé [VS Code](https://code.visualstudio.com/) . Git (pour cloner le dépôt du notebook) devrait être installé par défaut. * Un **compte Google** (pour s'authentifier avec Colab) * Recommandé : **Jupyter** extension (la plupart des configurations VS Code l'ont déjà) 2 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#installez-lextension-colab-dans-vs-code) Installez l'extension Colab dans VS Code 1. Ouvrez **Extensions** dans VS Code (`Ctrl+Shift+X` / `Cmd+Shift+X`) 2. Recherchez **« Colab »** et installez **l'extension Google Colab** extension ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FJ1f5BS4EK0ok5QLy6vz5%252Fcolab_img_1.png%3Falt%3Dmedia%26token%3D1faa7aac-c016-4c31-90ba-ccad655244e1&width=768&dpr=3&quality=100&sign=1c66f56a&sv=2) 3 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#ouvrir-un-notebook-unsloth) Ouvrir un notebook Unsloth 1. Clonez le dépôt [de notebooks Unsloth](https://github.com/unslothai/notebooks) : Copier git clone https://github.com/unslothai/notebooks cd notebooks/nb ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FwRL8pvgtLRpz5rcmS6vy%252FScreenshot%25202026-02-18%2520at%25206.29.16%25E2%2580%25AFPM.png%3Falt%3Dmedia%26token%3D7a72c2ec-6f94-4015-8355-27e76ba9b974&width=768&dpr=3&quality=100&sign=6104a846&sv=2) 1. Ouvrez le notebook souhaité. Unsloth prend en charge la plupart des modèles, y compris [l'embedding](https://unsloth.ai/docs/fr/notions-de-base/embedding-finetuning) , [TTS](https://unsloth.ai/docs/fr/notions-de-base/text-to-speech-tts-fine-tuning) . Par exemple, nous utiliserons Qwen3-4B RL : `nb/Qwen3_(4B)-GRPO.ipynb` ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FleZtIzyflSgi1o1QcZDU%252Fcolab_img_2.png%3Falt%3Dmedia%26token%3D805ebe8b-cb50-49a4-83d4-0daf96ad21c1&width=768&dpr=3&quality=100&sign=238de6cd&sv=2) 4 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#selectionnez-un-noyau-et-choisissez-colab) Sélectionnez un noyau et choisissez Colab Dans la barre d'outils du notebook, cliquez sur **Sélectionner le noyau**, puis choisissez **Colab** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FleGtAKSVehrlspPGvDy2%252Fcolab_img_3.png%3Falt%3Dmedia%26token%3Dba3d8f1d-9cbb-44ed-8e52-8d0e52280607&width=768&dpr=3&quality=100&sign=65818fa6&sv=2) 5 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#ajouter-un-nouveau-serveur-colab) Ajouter un nouveau serveur Colab Après avoir choisi **Colab**, vous verrez un menu déroulant avec des options de serveur. 1. Cliquez **\+ Ajouter un nouveau serveur Colab** 2. La première fois, une fenêtre de navigateur peut s'ouvrir pour l'authentification Google * Connectez-vous, accordez l'accès, puis revenez à VS Code ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FZcDsB2i0aUStIKEeFTvF%252Fcolab_img_4.png%3Falt%3Dmedia%26token%3Da7c3642f-db35-4215-933f-99c17ebc92a2&width=768&dpr=3&quality=100&sign=3adde5c1&sv=2) 6 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#selectionnez-le-gpu-et-nommez-le-serveur) Sélectionnez le GPU et nommez le serveur 1. Définissez **Accélérateur matériel** sur **GPU** 2. Choisissez un type de GPU (par exemple **T4**, si disponible) 3. Donnez un nom au serveur (n'importe quel nom de votre choix) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FLuK8oheihdCNn9MAZSlk%252Fcolab_img_5.png%3Falt%3Dmedia%26token%3De5ed396f-afe2-4c09-b1b8-120c7301793c&width=768&dpr=3&quality=100&sign=5e27c8e4&sv=2) Remarque : la disponibilité des GPU dépend de votre plan Colab et de la capacité actuelle. Si vous ne voyez pas d'options GPU, consultez le dépannage ci-dessous. 7 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#choisir-le-noyau-python) Choisir le noyau Python Une fois connecté au serveur Colab, sélectionnez le **noyau Python** qui apparaît pour ce runtime (généralement un noyau Python 3). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FBexUiXJ7LRFBXs95tX2S%252Fcolab_img_6.png%3Falt%3Dmedia%26token%3D391e343d-d2fa-4a77-8e3f-ef244adfcb4c&width=768&dpr=3&quality=100&sign=7466c2cb&sv=2) 8 #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#executer-le-notebook) Exécuter le notebook * Cliquez **Exécuter tout** dans la barre d'outils du notebook (ou exécutez les cellules de haut en bas) * Regardez les cellules de configuration installer les dépendances puis lancer le workflow Unsloth * Vous pouvez consulter nos guides dédiés [de fine-tuning](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) ou [d'apprentissage par renforcement](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) pour plus d'informations sur la manière de commencer précisément avec Unsloth. ### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#tutoriel-video) Tutoriel vidéo ### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#depannage) Dépannage #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#apres-la-deconnexion-du-serveur-colab-le-notebook-ne-sexecute-pas-sur-un-nouveau-serveur) Après la déconnexion du serveur Colab, le notebook ne s'exécute pas sur un nouveau serveur **Ce qui se passe :** Si le notebook reste ouvert pendant la déconnexion du serveur Colab, VS Code peut rester bloqué dans un mauvais état de noyau/runtime après la reconnexion. Problème lié sur [le dépôt GitHub](https://github.com/googlecolab/colab-vscode/issues/200) . **Correction :** Fermez complètement l'onglet du notebook et ouvrez de nouveau le notebook. #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#vous-ne-pouvez-pas-selectionner-un-gpu-seul-le-cpu-apparait) Vous ne pouvez pas sélectionner un GPU (seul le CPU apparaît) Causes possibles et solutions : * **Capacité du niveau gratuit de Colab :** Les GPU peuvent être temporairement indisponibles → réessayez plus tard. * **Pas réellement connecté à un runtime Colab :** vérifiez de nouveau **Sélectionner le noyau → Colab** et assurez-vous qu'un serveur Colab est actif. * **Restrictions de compte/région ou limites atteintes :** vous devrez peut-être attendre ou utiliser un autre compte Google / un autre plan. #### [](https://unsloth.ai/docs/fr/commencer/install/vs-code#tout-a-fonctionne-mais-les-paquets-ont-disparu-apres-la-reconnexion) Tout a fonctionné, mais les paquets ont “disparu” après la reconnexion Les runtimes Colab sont **éphémères**. Lorsqu'un serveur redémarre, vous devez généralement relancer les cellules de configuration/d'installation (souvent les premières cellules du notebook). Mis à jour il y a 5 mois Ce contenu vous a-t-il été utile ? * [Tutoriel VS Code et Colab :](https://unsloth.ai/docs/fr/commencer/install/vs-code#tutoriel-vs-code-et-colab) * [Tutoriel vidéo](https://unsloth.ai/docs/fr/commencer/install/vs-code#tutoriel-video) * [Dépannage](https://unsloth.ai/docs/fr/commencer/install/vs-code#depannage) Ce contenu vous a-t-il été utile ? --- # Installer Unsloth via Docker | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/commencer/install/docker.md) . Découvrez comment utiliser nos conteneurs Docker avec toutes les dépendances préinstallées pour une installation immédiate. Aucune configuration requise, il suffit de lancer et de commencer l’entraînement ! Image Docker Unsloth : [`**unsloth/unsloth**`](https://hub.docker.com/r/unsloth/unsloth) Unsloth Studio partage désormais le même cache que les notebooks et les scripts afin d’éviter les re-téléchargements inutiles. ### [](https://unsloth.ai/docs/fr/commencer/install/docker#demarrage-rapide) ⚡ Démarrage rapide 1 **Installez Docker et NVIDIA Container Toolkit.** Installez Docker via [Linux](https://docs.docker.com/engine/install/) ou [Desktop](https://docs.docker.com/desktop/) (autre). Puis installez [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#installation) : Copier export NVIDIA_CONTAINER_TOOLKIT_VERSION=1.17.8-1 sudo apt-get update && sudo apt-get install -y \ nvidia-container-toolkit=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ nvidia-container-toolkit-base=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ libnvidia-container-tools=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ libnvidia-container1=${NVIDIA_CONTAINER_TOOLKIT_VERSION} ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-41cae231ed4761f844ce9836e03b17aabd7c803c%252Fnvidia%2520toolkit.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=4eaf403&sv=2) 2 **Lancez le conteneur.** [`**unsloth/unsloth**`](https://hub.docker.com/r/unsloth/unsloth) est la seule image Docker d’Unsloth. Copier docker run -d -e JUPYTER_PASSWORD="mypassword" \ -p 8888:8888 -p 8000:8000 -p 2222:22 \ -v $(pwd)/work:/workspace/work \ --gpus all \ unsloth/unsloth ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-2b50d78c5d54eaf189c0a40d46c405585ea23082%252Fdocker%2520run.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=6b43716&sv=2) 3 **Accéder à Jupyter Lab** Rendez-vous sur [http://localhost:8888](http://localhost:8888/) et ouvrez Unsloth. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-828df0a668fd94025c1193c24a7f09c1d58dcbd8%252Fjupyter.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=3b6f4e7b&sv=2) Accédez aux onglets `unsloth-notebooks` pour voir les notebooks Unsloth. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-e7a3f620a3ec5bff335632ff9b0cb422f76528a1%252FScreenshot_from_2025-09-30_21-38-15.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=4d64108f&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-531882c33eb96dec24e2d7673471d6a3928a3951%252FScreenshot_from_2025-09-30_21-39-41.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=422ce13a&sv=2) 4 **Commencez l’entraînement avec Unsloth** Si vous débutez, suivez notre guide pas à pas [Guide de fine-tuning](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide) , [Guide RL](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide) ou enregistrez/copie simplement l’un de nos [notebooks](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-665f900b008991ddcd8fdabb773b292de3c41e72%252FScreenshot_from_2025-09-30_21-40-29.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=7d746efd&sv=2) #### [](https://unsloth.ai/docs/fr/commencer/install/docker#structure-du-conteneur) 📂 Structure du conteneur * `/workspace/work/` — Votre répertoire de travail monté * `/workspace/unsloth-notebooks/` — Exemples de notebooks de fine-tuning * `/home/unsloth/` — Répertoire personnel de l’utilisateur ### [](https://unsloth.ai/docs/fr/commencer/install/docker#exemple-dutilisation) 📖 Exemple d’utilisation #### [](https://unsloth.ai/docs/fr/commencer/install/docker#exemple-complet) Exemple complet Copier docker run -d -e JUPYTER_PORT=8000 \ -e JUPYTER_PASSWORD="mypassword" \ -e "SSH_KEY=$(cat ~/.ssh/container_key.pub)" \ -e USER_PASSWORD="unsloth2024" \ -p 8000:8000 -p 2222:22 \ -v $(pwd)/work:/workspace/work \ --gpus all \ unsloth/unsloth #### [](https://unsloth.ai/docs/fr/commencer/install/docker#configuration-de-la-cle-ssh) Configuration de la clé SSH Si vous n’avez pas de paire de clés SSH : ### [](https://unsloth.ai/docs/fr/commencer/install/docker#pourquoi-les-conteneurs-unsloth) 🦥Pourquoi les conteneurs Unsloth ? * **Fiable**: Environnement sélectionné avec des versions de paquets stables et maintenues. Seulement 7 Go compressés (contre 10–11 Go ailleurs) * **Prêt à l’emploi**: Notebooks préinstallés dans `/workspace/unsloth-notebooks/` * **Sécurisé**: S’exécute en toute sécurité en tant qu’utilisateur non-root * **Universel**: Compatible avec tous les modèles basés sur des transformeurs (TTS, BERT, etc.) ### [](https://unsloth.ai/docs/fr/commencer/install/docker#unsloth-ne-detecte-pas-ou-nutilise-pas-mon-gpu) **Unsloth ne détecte pas ou n’utilise pas mon GPU** Si le modèle n’utilise pas votre GPU spécifiquement pour Docker, essayez : En tirant manuellement la dernière image : * Démarrez le conteneur avec l’accès GPU : * `docker run`: `--gpus all` * Docker Compose : `capabilities: [gpu]` * Sous Linux, assurez-vous que NVIDIA Container Toolkit est installé. * Sous Windows : * Vérifiez que `nvcc --version` correspond à la version CUDA affichée dans `nvidia-smi` * Suivez : [https://docs.docker.com/desktop/features/gpu/](https://docs.docker.com/desktop/features/gpu/) ### [](https://unsloth.ai/docs/fr/commencer/install/docker#parametres-avances) ⚙️ Paramètres avancés Variable Description Valeur par défaut `JUPYTER_PASSWORD` Mot de passe de Jupyter Lab `unsloth` `JUPYTER_PORT` Port de Jupyter Lab à l’intérieur du conteneur `8888` `SSH_KEY` Clé publique SSH pour l’authentification `Aucune` `USER_PASSWORD` Mot de passe pour `unsloth` utilisateur (sudo) `unsloth` * Jupyter Lab : `-p 8000:8888` * Accès SSH : `-p 2222:22` **Important**: Utilisez des montages de volumes pour conserver votre travail entre les exécutions du conteneur. ### [](https://unsloth.ai/docs/fr/commencer/install/docker#notes-de-securite) **🔒 Notes de sécurité** * Le conteneur s’exécute par défaut en tant que non-root `unsloth` utilisateur * Utilisez `USER_PASSWORD` pour les opérations sudo à l’intérieur du conteneur * L’accès SSH nécessite une authentification par clé publique [PrécédentAMD](https://unsloth.ai/docs/fr/commencer/install/amd) [SuivantUpdating](https://unsloth.ai/docs/fr/commencer/install/updating) Mis à jour il y a 3 mois Ce contenu vous a-t-il été utile ? * [⚡ Démarrage rapide](https://unsloth.ai/docs/fr/commencer/install/docker#demarrage-rapide) * [📖 Exemple d’utilisation](https://unsloth.ai/docs/fr/commencer/install/docker#exemple-dutilisation) * [🦥Pourquoi les conteneurs Unsloth ?](https://unsloth.ai/docs/fr/commencer/install/docker#pourquoi-les-conteneurs-unsloth) * [Unsloth ne détecte pas ou n’utilise pas mon GPU](https://unsloth.ai/docs/fr/commencer/install/docker#unsloth-ne-detecte-pas-ou-nutilise-pas-mon-gpu) * [⚙️ Paramètres avancés](https://unsloth.ai/docs/fr/commencer/install/docker#parametres-avances) * [🔒 Notes de sécurité](https://unsloth.ai/docs/fr/commencer/install/docker#notes-de-securite) Ce contenu vous a-t-il été utile ? Copier # Générer une nouvelle paire de clés ssh-keygen -t rsa -b 4096 -f ~/.ssh/container_key # Utiliser la clé publique dans docker run -e "SSH_KEY=$(cat ~/.ssh/container_key.pub)" # Se connecter via SSH ssh -i ~/.ssh/container_key -p 2222 unsloth@localhost Copier docker pull unsloth/unsloth:latest Copier # Générer une paire de clés SSH ssh-keygen -t rsa -b 4096 -f ~/.ssh/container_key # Se connecter au conteneur ssh -i ~/.ssh/container_key -p 2222 unsloth@localhost Copier -p : Copier -v : Copier docker run -d -e JUPYTER_PORT=8000 \ -e JUPYTER_PASSWORD="mypassword" \ -e "SSH_KEY=$(cat ~/.ssh/container_key.pub)" \ -e USER_PASSWORD="unsloth2024" \ -p 8000:8000 -p 2222:22 \ -v $(pwd)/work:/workspace/work \ --gpus all \ unsloth/unsloth --- # Hugging Face Hub, XET debugging | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) . #### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging#downloads-are-stuck-at-90-to-99) Downloads are stuck at 90% to 99% ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FAupnur7cpMvilAleA1Gp%252Fimage.png%3Falt%3Dmedia%26token%3D146a3d2e-6c2c-438e-9bf7-6265bc6d6cb4&width=768&dpr=3&quality=100&sign=2c7861fd&sv=2) If you see downloads via `hf download unsloth/*` get stuck at 90% or 99% of progress for quite some time, cancel the current run, and try adding using below commands: #### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging#rate-limited-or-429-too-many-requests) Rate limited or 429 Too Many Requests? Try using `snapshot_download` instead, then import Unsloth which will set the correct Hugging Face variables for you: Or maybe try getting a Hugging Face token first via [https://huggingface.co/settings/tokens](https://huggingface.co/settings/tokens) [PreviousTroubleshooting & FAQs](https://unsloth.ai/docs/basics/troubleshooting-and-faqs) [NextChat Templates](https://unsloth.ai/docs/basics/chat-templates) Last updated 5 months ago Was this helpful? Was this helpful? Copy pip install -U huggingface_hub HF_HOME=".cache_new/huggingface" \ HF_XET_CACHE=".cache_new/huggingface/xet" \ HF_HUB_CACHE=".cache_new/huggingface/hub" \ HF_XET_HIGH_PERFORMANCE=1 \ HF_XET_CHUNK_CACHE_SIZE_BYTES=0 \ HF_XET_RECONSTRUCT_WRITE_SEQUENTIALLY=0 \ HF_XET_NUM_CONCURRENT_RANGE_GETS=64 \ hf download unsloth/Qwen3-Coder-Next-GGUF \ --local-dir unsloth/Qwen3-Coder-Next-GGUF \ --include "*UD-Q6_K_XL*" Copy import unsloth import os os.environ["HF_HOME"] = ".cache_new/huggingface" os.environ["HF_XET_CACHE"] = ".cache_new/huggingface/xet" os.environ["HF_HUB_CACHE"] = ".cache_new/huggingface/hub" from huggingface_hub import snapshot_download snapshot_download( repo_id = "unsloth/Qwen3-Coder-Next-GGUF", local_dir = "unsloth/Qwen3-Coder-Next-GGUF", allow_patterns = ["*UD-Q6_K_XL*"], ) Copy pip install -U huggingface_hub HF_HOME=".cache_new/huggingface" \ HF_XET_CACHE=".cache_new/huggingface/xet" \ HF_HUB_CACHE=".cache_new/huggingface/hub" \ HF_XET_HIGH_PERFORMANCE=1 \ HF_XET_CHUNK_CACHE_SIZE_BYTES=0 \ HF_XET_RECONSTRUCT_WRITE_SEQUENTIALLY=0 \ HF_XET_NUM_CONCURRENT_RANGE_GETS=64 \ hf download unsloth/Qwen3-Coder-Next-GGUF \ --local-dir unsloth/Qwen3-Coder-Next-GGUF \ --include "*UD-Q6_K_XL*" \ --token "hf_ADD_YOUR_HUGGING_FACE_TOKEN_HERE" --- # Continued Pretraining | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/continued-pretraining.md) . * The [text completion notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral_(7B)-Text_Completion.ipynb) is for continued pretraining/raw text. * The [continued pretraining notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral_v0.3_(7B)-CPT.ipynb) is for learning another language. You can read more about continued pretraining and our release in our [blog post](https://unsloth.ai/blog/contpretraining) . [](https://unsloth.ai/docs/basics/continued-pretraining#what-is-continued-pretraining) What is Continued Pretraining? -------------------------------------------------------------------------------------------------------------------------- Continued or continual pretraining (CPT) is necessary to “steer” the language model to understand new domains of knowledge, or out of distribution domains. Base models like Llama-3 8b or Mistral 7b are first pretrained on gigantic datasets of trillions of tokens (Llama-3 for e.g. is 15 trillion). But sometimes these models have not been well trained on other languages, or text specific domains, like law, medicine or other areas. So continued pretraining (CPT) is necessary to make the language model learn new tokens or datasets. [](https://unsloth.ai/docs/basics/continued-pretraining#advanced-features) Advanced Features: -------------------------------------------------------------------------------------------------- ### [](https://unsloth.ai/docs/basics/continued-pretraining#loading-lora-adapters-for-continued-finetuning) Loading LoRA adapters for continued finetuning If you saved a LoRA adapter through Unsloth, you can also continue training using your LoRA weights. The optimizer state will be reset as well. To load even optimizer states to continue finetuning, see the next section. Copy from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name = "LORA_MODEL_NAME", max_seq_length = max_seq_length, dtype = dtype, load_in_4bit = load_in_4bit, ) trainer = Trainer(...) trainer.train() ### [](https://unsloth.ai/docs/basics/continued-pretraining#continued-pretraining-and-finetuning-the-lm_head-and-embed_tokens-matrices) Continued Pretraining & Finetuning the `lm_head` and `embed_tokens` matrices Add `lm_head` and `embed_tokens`. For Colab, sometimes you will go out of memory for Llama-3 8b. If so, just add `lm_head`. Then use 2 different learning rates - a 2-10x smaller one for the `lm_head` or `embed_tokens` like so: [PreviousChat Templates](https://unsloth.ai/docs/basics/chat-templates) [NextUnsloth Start](https://unsloth.ai/docs/integrations/unsloth-start) Last updated 8 months ago Was this helpful? * [What is Continued Pretraining?](https://unsloth.ai/docs/basics/continued-pretraining#what-is-continued-pretraining) * [Advanced Features:](https://unsloth.ai/docs/basics/continued-pretraining#advanced-features) * [Loading LoRA adapters for continued finetuning](https://unsloth.ai/docs/basics/continued-pretraining#loading-lora-adapters-for-continued-finetuning) * [Continued Pretraining & Finetuning the lm\_head and embed\_tokens matrices](https://unsloth.ai/docs/basics/continued-pretraining#continued-pretraining-and-finetuning-the-lm_head-and-embed_tokens-matrices) Was this helpful? Copy model = FastLanguageModel.get_peft_model( model, r = 16, target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",\ "gate_proj", "up_proj", "down_proj",\ "lm_head", "embed_tokens",], lora_alpha = 16, ) Copy from unsloth import UnslothTrainer, UnslothTrainingArguments trainer = UnslothTrainer( .... args = UnslothTrainingArguments( .... learning_rate = 5e-5, embedding_learning_rate = 5e-6, # 2-10x smaller than learning_rate ), ) --- # Unsloth Environment Flags | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/unsloth-environment-flags.md) . Environment variable Purpose `os.environ["UNSLOTH_RETURN_LOGITS"] = "1"` Forcibly returns logits - useful for evaluation if logits are needed. `os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"` Disables auto compiler. Could be useful to debug incorrect finetune results. `os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"` Disables fast generation for generic models. `os.environ["UNSLOTH_ENABLE_LOGGING"] = "1"` Enables auto compiler logging - useful to see which functions are compiled or not. `os.environ["UNSLOTH_FORCE_FLOAT32"] = "1"` On float16 machines, use float32 and not float16 mixed precision. Useful for Gemma 3. `os.environ["UNSLOTH_STUDIO_DISABLED"] = "1"` Disables extra features. `os.environ["UNSLOTH_COMPILE_DEBUG"] = "1"` Turns on extremely verbose `torch.compile`logs. `os.environ["UNSLOTH_COMPILE_MAXIMUM"] = "0"` Enables maximum `torch.compile`optimizations - not recommended. `os.environ["UNSLOTH_COMPILE_IGNORE_ERRORS"] = "1"` Can turn this off to enable fullgraph parsing. `os.environ["UNSLOTH_FULLGRAPH"] = "0"` Enable `torch.compile` fullgraph mode `os.environ["UNSLOTH_DISABLE_AUTO_UPDATES"] = "1"` Forces no updates to `unsloth-zoo` Another possibility is maybe the model uploads we uploaded are corrupted, but unlikely. Try the following: Copy model, tokenizer = FastVisionModel.from_pretrained( "Qwen/Qwen2-VL-7B-Instruct", use_exact_model_name = True, ) Last updated 2 months ago Was this helpful? Was this helpful? --- # Multi-GPU Fine-tuning with Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth.md) . Unsloth currently supports multi-GPU setups through libraries like Accelerate and DeepSpeed. This means you can already leverage parallelism methods such as **FSDP** and **DDP** with Unsloth. #### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth#see-our-new-distributed-data-parallel-ddp-multi-gpu-guide-here) **See our new Distributed Data Parallel** [**(DDP) multi-GPU Guide here**](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp) **.** We know that the process can be complex and requires manual setup. We’re working hard to make multi-GPU support much simpler and more user-friendly, and we’ll be announcing official multi-GPU support for Unsloth soon. For now, you can use our [Magistral-2509 Kaggle notebook](https://unsloth.ai/docs/models/tutorials/magistral-how-to-run-and-fine-tune#fine-tuning-magistral-with-unsloth) as an example which utilizes multi-GPU Unsloth to fit the 24B parameter model or our [DDP guide](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp) . **In the meantime**, to enable multi GPU for DDP, do the following: 1. Create your training script as `train.py` (or similar). For example, you can use one of our [training scripts](https://github.com/unslothai/notebooks/tree/main/python_scripts) created from our various notebooks! 2. Run `accelerate launch train.py` or `torchrun --nproc_per_node N_GPUS train.py` where `N_GPUS` is the number of GPUs you have. #### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth#pipeline-model-splitting-loading) **Pipeline / model splitting loading** If you do not have enough VRAM for 1 GPU to load say Llama 70B, no worries - we will split the model for you on each GPU! To enable this, use the `device_map = "balanced"` flag: Copy from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( "unsloth/Llama-3.3-70B-Instruct", load_in_4bit = True, device_map = "balanced", ) **Stay tuned for our official announcement!** For more details, check out our ongoing [Pull Request](https://github.com/unslothai/unsloth/issues/2435) discussing multi-GPU support. [PreviousMCP Server](https://unsloth.ai/docs/basics/mcp) [NextDistributed Data Parallel (DDP)](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp) Last updated 4 months ago Was this helpful? Was this helpful? --- # Unsloth Benchmarks | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/unsloth-benchmarks.md) . * For more detailed benchmarks, read our [Llama 3.3 Blog](https://unsloth.ai/blog/llama3-3) . * Benchmarking of Unsloth was also conducted by [🤗Hugging Face](https://huggingface.co/blog/unsloth-trl) . If your speed seems slower at first, it’s likely because `torch.compile` typically takes ~5 minutes (or longer) to warm up and finish compiling. Make sure you measure throughput **after** it’s fully loaded as over longer runs, Unsloth should be much faster. Tested on H100 and [Blackwell](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth) GPUs. We tested using the Alpaca Dataset, a batch size of 2, gradient accumulation steps of 4, rank = 32, and applied QLoRA on all linear layers (q, k, v, o, gate, up, down): Model VRAM 🦥Unsloth speed 🦥VRAM reduction 🦥Longer context 😊Hugging Face + FA2 Llama 3.3 (70B) 80GB 2x \>75% 13x longer 1x Llama 3.1 (8B) 80GB 2x \>70% 12x longer 1x [](https://unsloth.ai/docs/basics/unsloth-benchmarks#context-length-benchmarks) Context length benchmarks -------------------------------------------------------------------------------------------------------------- The more data you have, the less VRAM Unsloth uses due to our [gradient checkpointing](https://unsloth.ai/blog/long-context) algorithm + Apple's CCE algorithm! ### [](https://unsloth.ai/docs/basics/unsloth-benchmarks#llama-3.1-8b-max.-context-length) **Llama 3.1 (8B) max. context length** We tested Llama 3.1 (8B) Instruct and did 4bit QLoRA on all linear layers (Q, K, V, O, gate, up and down) with rank = 32 with a batch size of 1. We padded all sequences to a certain maximum sequence length to mimic long context finetuning workloads. GPU VRAM 🦥Unsloth context length Hugging Face + FA2 8 GB 2,972 OOM 12 GB 21,848 932 16 GB 40,724 2,551 24 GB 78,475 5,789 40 GB 153,977 12,264 48 GB 191,728 15,502 80 GB 342,733 28,454 ### [](https://unsloth.ai/docs/basics/unsloth-benchmarks#llama-3.3-70b-max.-context-length) **Llama 3.3 (70B) max. context length** We tested Llama 3.3 (70B) Instruct on a 80GB A100 and did 4bit QLoRA on all linear layers (Q, K, V, O, gate, up and down) with rank = 32 with a batch size of 1. We padded all sequences to a certain maximum sequence length to mimic long context finetuning workloads. GPU VRAM 🦥Unsloth context length Hugging Face + FA2 48 GB 12,106 OOM 80 GB 89,389 6,916 Last updated 2 months ago Was this helpful? * [Context length benchmarks](https://unsloth.ai/docs/basics/unsloth-benchmarks#context-length-benchmarks) * [Llama 3.1 (8B) max. context length](https://unsloth.ai/docs/basics/unsloth-benchmarks#llama-3.1-8b-max.-context-length) * [Llama 3.3 (70B) max. context length](https://unsloth.ai/docs/basics/unsloth-benchmarks#llama-3.3-70b-max.-context-length) Was this helpful? --- # Finetuning from Last Checkpoint | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint.md) . You must edit the `Trainer` first to add `save_strategy` and `save_steps`. Below saves a checkpoint every 50 steps to the folder `outputs`. Copy trainer = SFTTrainer( .... args = TrainingArguments( .... output_dir = "outputs", save_strategy = "steps", save_steps = 50, ), ) Then in the trainer do: Copy trainer_stats = trainer.train(resume_from_checkpoint = True) Which will start from the latest checkpoint and continue training. ### [](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint#wandb-integration) Wandb Integration Copy # Install library !pip install wandb --upgrade # Setting up Wandb !wandb login import os os.environ["WANDB_PROJECT"] = "" os.environ["WANDB_LOG_MODEL"] = "checkpoint" Then in `TrainingArguments()` set To train the model, do `trainer.train()`; to resume training, do [](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint#how-do-i-do-early-stopping) ❓How do I do Early Stopping? ------------------------------------------------------------------------------------------------------------------------------- If you want to stop or pause the finetuning / training run since the evaluation loss is not decreasing, then you can use early stopping which stops the training process. Use `EarlyStoppingCallback`. As usual, set up your trainer and your evaluation dataset. The below is used to stop the training run if the `eval_loss` (the evaluation loss) is not decreasing after 3 steps or so. We then add the callback which can also be customized: Then train the model as usual via `trainer.train() .` Last updated 2 months ago Was this helpful? * [Wandb Integration](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint#wandb-integration) * [❓How do I do Early Stopping?](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint#how-do-i-do-early-stopping) Was this helpful? Copy report_to = "wandb", logging_steps = 1, # Change if needed save_steps = 100 # Change if needed run_name = "" # (Optional) Copy import wandb run = wandb.init() artifact = run.use_artifact('//', type='model') artifact_dir = artifact.download() trainer.train(resume_from_checkpoint=artifact_dir) Copy from trl import SFTConfig, SFTTrainer trainer = SFTTrainer( args = SFTConfig( fp16_full_eval = True, per_device_eval_batch_size = 2, eval_accumulation_steps = 4, output_dir = "training_checkpoints", # location of saved checkpoints for early stopping save_strategy = "steps", # save model every N steps save_steps = 10, # how many steps until we save the model save_total_limit = 3, # keep only 3 saved checkpoints to save disk space eval_strategy = "steps", # evaluate every N steps eval_steps = 10, # how many steps until we do evaluation load_best_model_at_end = True, # MUST USE for early stopping metric_for_best_model = "eval_loss", # metric we want to early stop on greater_is_better = False, # the lower the eval loss, the better ), model = model, tokenizer = tokenizer, train_dataset = new_dataset["train"], eval_dataset = new_dataset["test"], ) Copy from transformers import EarlyStoppingCallback early_stopping_callback = EarlyStoppingCallback( early_stopping_patience = 3, # How many steps we will wait if the eval loss doesn't decrease # For example the loss might increase, but decrease after 3 steps early_stopping_threshold = 0.0, # Can set higher - sets how much loss should decrease by until # we consider early stopping. For eg 0.01 means if loss was # 0.02 then 0.01, we consider to early stop the run. ) trainer.add_callback(early_stopping_callback) --- # GPU Mode - Reinforcement Learning Mini Conference 2026 | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/gpu-mode-conference.md) . Unsloth GitHub for efficient GRPO / GSPO with lots of custom kernel work: [https://github.com/unslothai/unsloth](https://github.com/unslothai/unsloth) Here are our attached PDF slides: 12MB [RL Mini Conference \_\_ Unsloth.pdf](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F6LJafOec6iUGTtyLr2c5%2FRL%20Mini%20Conference%20__%20Unsloth.pdf?alt=media&token=ac3eeb48-8d54-424b-b8a9-3ba02f13efe2) PDF Download[Open](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F6LJafOec6iUGTtyLr2c5%2FRL%20Mini%20Conference%20__%20Unsloth.pdf?alt=media&token=ac3eeb48-8d54-424b-b8a9-3ba02f13efe2) ### [](https://unsloth.ai/docs/blog/gpu-mode-conference#reinforcement-learning-resources) Reinforcement Learning Resources For step-by-step guides, beginner or advanced tutorials, you can refer to our docs for: [💡Reinforcement Learning](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide) [RL Reward Hacking](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/rl-reward-hacking) [🧩Advanced RL Docs](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation) [⁉️FP16 vs BF16 for RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl) [🎱FP8 RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning) [👁️‍🗨️Vision RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl) ### [](https://unsloth.ai/docs/blog/gpu-mode-conference#live-video-recording) Live Video Recording Last updated 6 months ago Was this helpful? * [Reinforcement Learning Resources](https://unsloth.ai/docs/blog/gpu-mode-conference#reinforcement-learning-resources) * [Live Video Recording](https://unsloth.ai/docs/blog/gpu-mode-conference#live-video-recording) Was this helpful? --- # Quantization-Aware Training (QAT) | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/quantization-aware-training-qat.md) . In collaboration with PyTorch, we're introducing QAT (Quantization-Aware Training) in Unsloth to enable **trainable quantization** that recovers as much accuracy as possible. This results in significantly better model quality compared to standard 4-bit naive quantization. QAT can recover up to **70% of the lost accuracy** and achieve a **1–3%** model performance improvement on benchmarks such as GPQA and MMLU Pro. > **Try QAT with our free** [**Qwen3 (4B) notebook**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(4B)_Instruct-QAT.ipynb) ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantization) 📚Quantization Naively quantizing a model is called **post-training quantization** (PTQ). For example, assume we want to quantize to 8bit integers: 1. Find `max(abs(W))` 2. Find `a = 127/max(abs(W))` where a is int8's maximum range which is 127 3. Quantize via `qW = int8(round(W * a))` ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-f3e1cee8e4047dcbbbace7548694ad63af9869de%252Fquant-freeze.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=27520418&sv=2) Dequantizing back to 16bits simply does the reverse operation by `float16(qW) / a` . Post-training quantization (PTQ) can greatly reduce storage and inference costs, but quite often degrades accuracy when representing high-precision values with fewer bits - especially at 4-bit or lower. One way to solve this to utilize our [**dynamic GGUF quants**](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs) , which uses a calibration dataset to change the quantization procedure to allocate more importance to important weights. The other way is to make **quantization smarter, by making it trainable or learnable**! ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#smarter-quantization) 🔥Smarter Quantization ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-1f6260ef5c041ada2f8b1fb4c6aad114f61061d4%252F4bit_QAT_recovery_sideways_clipped75_bigtext_all%281%29.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=db52d2bd&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-ad1ac9d29482ea07cbabb6efa18a0d1f06b297e9%252FQLoRA_QAT_Accuracy_Boosts_v7_bigaxes_nogrid_600dpi.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=c4c4a3d3&sv=2) To enable smarter quantization, we collaborated with the [TorchAO](https://github.com/pytorch/ao) team to add **Quantization-Aware Training (QAT)** directly inside of Unsloth - so now you can fine-tune models in Unsloth and then export them to 4-bit QAT format directly with accuracy improvements! In fact, **QAT recovers 66.9%** of Gemma3-4B on GPQA, and increasing the raw accuracy by +1.0%. Gemma3-12B on BBH recovers 45.5%, and **increased the raw accuracy by +2.1%**. QAT has no extra overhead during inference, and uses the same disk and memory usage as normal naive quantization! So you get all the benefits of low-bit quantization, but with much increased accuracy! ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantization-aware-training) 🔍Quantization-Aware Training QAT simulates the true quantization procedure by "**fake quantizing**" weights and optionally activations during training, which typically means rounding high precision values to quantized ones (while staying in high precision dtype, e.g. bfloat16) and then immediately dequantizing them. TorchAO enables QAT by first (1) inserting fake quantize operations into linear layers, and (2) transforms the fake quantize operations to actual quantize and dequantize operations after training to make it inference ready. Step 1 enables us to train a more accurate quantization representation. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-3d990e2bf19ef1aa7e65a8dd07e4b71cf8882a2a%252Fqat_diagram.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=899f09a1&sv=2) ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#qat--lora-finetuning) ✨QAT + LoRA finetuning QAT in Unsloth can additionally be combined with LoRA fine-tuning to enable the benefits of both worlds: significantly reducing storage and compute requirements during training while mitigating quantization degradation! We support multiple methods via `qat_scheme` including `fp8-int4`, `fp8-fp8`, `int8-int4`, `int4` . We also plan to add custom definitions for QAT in a follow up release! ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#exporting-qat-models) 🫖Exporting QAT models After fine-tuning in Unsloth, you can call `model.save_pretrained_torchao` to save your trained model using TorchAO’s PTQ format. You can also upload these to the HuggingFace hub! We support any config, and we plan to make text based methods as well, and to make the process more simpler for everyone! But first, we have to prepare the QAT model for the final conversion step via: And now we can select which QAT style you want: You can then run the merged QAT lower precision model in vLLM, Unsloth and other systems for inference! These are all in the [Qwen3-4B QAT Colab notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(4B)_Instruct-QAT.ipynb) we have as well! ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantizing-models-without-training) 🫖Quantizing models without training You can also call `model.save_pretrained_torchao` directly without doing any QAT as well! This is simply PTQ or native quantization. For example, saving to Dynamic float8 format is below: ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#executorch-qat-for-mobile-deployment) 📱ExecuTorch - QAT for mobile deployment With Unsloth and TorchAO’s QAT support, you can also fine-tune a model in Unsloth and seamlessly export it to [ExecuTorch](https://github.com/pytorch/executorch) (PyTorch’s solution for on-device inference) and deploy it directly on mobile. See an example in action [here](https://huggingface.co/metascroy/Qwen3-4B-int8-int4-unsloth) with more detailed workflows on the way! **Announcement coming soon!** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-53631bae5588644d2c64cec18f371f0a7e2688c6%252Fswiftpm_xcode.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=d725d0ee&sv=2) ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#how-to-enable-qat) 🌻How to enable QAT Update Unsloth to the latest version, and also install the latest TorchAO! Then **try QAT with our free** [**Qwen3 (4B) notebook**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(4B)_Instruct-QAT.ipynb) ### [](https://unsloth.ai/docs/blog/quantization-aware-training-qat#acknowledgements) 💁Acknowledgements Huge thanks to the entire PyTorch and TorchAO team for their help and collaboration! Extreme thanks to Andrew Or, Jerry Zhang, Supriya Rao, Scott Roy and Mergen Nachin for helping on many discussions on QAT, and on helping to integrate it into Unsloth! Also thanks to the Executorch team as well! [Previous500K Context Training](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning) [NextDGX Station](https://unsloth.ai/docs/blog/dgx-station) Last updated 7 months ago Was this helpful? * [📚Quantization](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantization) * [🔥Smarter Quantization](https://unsloth.ai/docs/blog/quantization-aware-training-qat#smarter-quantization) * [🔍Quantization-Aware Training](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantization-aware-training) * [✨QAT + LoRA finetuning](https://unsloth.ai/docs/blog/quantization-aware-training-qat#qat--lora-finetuning) * [🫖Exporting QAT models](https://unsloth.ai/docs/blog/quantization-aware-training-qat#exporting-qat-models) * [🫖Quantizing models without training](https://unsloth.ai/docs/blog/quantization-aware-training-qat#quantizing-models-without-training) * [📱ExecuTorch - QAT for mobile deployment](https://unsloth.ai/docs/blog/quantization-aware-training-qat#executorch-qat-for-mobile-deployment) * [🌻How to enable QAT](https://unsloth.ai/docs/blog/quantization-aware-training-qat#how-to-enable-qat) * [💁Acknowledgements](https://unsloth.ai/docs/blog/quantization-aware-training-qat#acknowledgements) Was this helpful? Copy from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name = "unsloth/Qwen3-4B-Instruct-2507", max_seq_length = 2048, load_in_16bit = True, ) model = FastLanguageModel.get_peft_model( model, r = 16, target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",\ "gate_proj", "up_proj", "down_proj",], lora_alpha = 32, # We support fp8-int4, fp8-fp8, int8-int4, int4 qat_scheme = "int4", ) Copy from torchao.quantization import quantize_ from torchao.quantization.qat import QATConfig quantize_(model, QATConfig(step = "convert")) Copy # Use the exact same config as QAT (convenient function) model.save_pretrained_torchao( model, "tokenizer", torchao_config = model._torchao_config.base_config, ) # Int4 QAT from torchao.quantization import Int4WeightOnlyConfig model.save_pretrained_torchao( model, "tokenizer", torchao_config = Int4WeightOnlyConfig(), ) # Int8 QAT from torchao.quantization import Int8DynamicActivationInt8WeightConfig model.save_pretrained_torchao( model, "tokenizer", torchao_config = Int8DynamicActivationInt8WeightConfig(), ) Copy # Float8 from torchao.quantization import PerRow from torchao.quantization import Float8DynamicActivationFloat8WeightConfig torchao_config = Float8DynamicActivationFloat8WeightConfig(granularity = PerRow()) model.save_pretrained_torchao(torchao_config = torchao_config) Copy pip install --upgrade --no-cache-dir --force-reinstall unsloth unsloth_zoo pip install torchao==0.14.0 fbgemm-gpu-genai==1.3.0 --- # How to Connect OpenRouter to Unsloth: API Key & Model Setup | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/openrouter.md) . This guide explains how to connect **OpenRouter to** [**Unsloth**](https://github.com/unslothai/unsloth) so you can access hosted AI models from providers like **OpenAI, Anthropic,** and **Google** through an open-source local UI chat interface. You’ll learn how to create an OpenRouter API key, add OpenRouter as a provider in Unsloth, load or manually enter model IDs, and enable external models for chat. Once a single API key is connected, OpenRouter models in Unsloth can provide advanced features such as thinking, web search, tool calling, code execution, and customizable generation settings directly from the chat page. ### [](https://unsloth.ai/docs/integrations/connections/openrouter#setup) Setup 1 #### [](https://unsloth.ai/docs/integrations/connections/openrouter#create-an-openrouter-api-key) Create an OpenRouter API key Sign in to your OpenRouter account. Create an API key from the [OpenRouter dashboard](https://openrouter.ai/settings/keys) : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FumQXaeZGAQ8f3hMQSQQY%252Fimage.png%3Falt%3Dmedia%26token%3D51ec273d-4271-4683-8cf2-0fb11bccb1ba&width=768&dpr=3&quality=100&sign=92e56127&sv=2) Copy the key. You will paste it into Unsloth in the next step. When creating the key, you can optionally set a credit limit or expiration date. 2 #### [](https://unsloth.ai/docs/integrations/connections/openrouter#connect-openrouter-to-unsloth) Connect OpenRouter to Unsloth Open **Settings → Connections**, then click **Add Connected**. Select **OpenRouter**, then enter your connection details. Enter your OpenRouter details: * **API key:** paste your OpenRouter API key ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FB9fDpRnYfDjhCI5yDoon%252Fimage.png%3Falt%3Dmedia%26token%3D97ecc7e3-a3fe-4cd8-a5c6-738c9de7fbf3&width=768&dpr=3&quality=100&sign=468b6ae4&sv=2) * **Model IDs:** click **Load Models**, or enter model IDs manually ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F59F88IYS0BzteOqcavUj%252Fimage.png%3Falt%3Dmedia%26token%3D72632261-b3cd-486c-b4dd-09b56fa3a58a&width=768&dpr=3&quality=100&sign=cbb6fe5a&sv=2) Finally, click **Add Connection**. 3 #### [](https://unsloth.ai/docs/integrations/connections/openrouter#ready-to-chat) Ready to Chat After saving the connection, select an OpenRouter model under **Connected** in the model dropdown. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FFFvSXoLCnvydwmTgS7Ru%252Fimage.png%3Falt%3Dmedia%26token%3D79ca30bb-e049-4106-8b98-0871f79dfba5&width=768&dpr=3&quality=100&sign=d499db5f&sv=2) OpenRouter models can expose different controls depending on the upstream model, including web search, thinking, tool-calling, and generation settings. ### [](https://unsloth.ai/docs/integrations/connections/openrouter#model-selection) Model Selection OpenRouter provides access to many models from different providers. If **Load Models** does not return the models you want to select, enter the model IDs you want enabled. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FQM4eHEcNpWQazKtyS06X%252Fimage.png%3Falt%3Dmedia%26token%3D8774a4a0-7d00-4888-9e03-72c0e8641533&width=768&dpr=3&quality=100&sign=127c993a&sv=2) Example model IDs: ### [](https://unsloth.ai/docs/integrations/connections/openrouter#troubleshooting) Troubleshooting If OpenRouter fails to connect, check that the API key is valid and belongs to the correct OpenRouter account. If a model does not appear after clicking **Load Models**, it may not be available for your account or region. You can enter the model ID manually or choose another model. [PreviousOllama](https://unsloth.ai/docs/integrations/connections/ollama) [NextHermes Agent](https://unsloth.ai/docs/integrations/hermes-agent) Last updated 1 month ago Was this helpful? * [Setup](https://unsloth.ai/docs/integrations/connections/openrouter#setup) * [Model Selection](https://unsloth.ai/docs/integrations/connections/openrouter#model-selection) * [Troubleshooting](https://unsloth.ai/docs/integrations/connections/openrouter#troubleshooting) Was this helpful? Copy openai/gpt-5.5 anthropic/claude-sonnet-4.6 google/gemini-3-pro --- # 500K Context Length Fine-tuning | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning.md) . We’re introducing new algorithms in Unsloth that push the limits of long-context training for **any LLM and VLM**. Training LLMs like gpt-oss-20b can now reach **500K+ context lengths** on a single 80GB H100 GPU, compared to 80K previously with no accuracy degradation. You can reach >**750K context windows** on a B200 192GB GPU. > **Try 500K-context gpt-oss-20b fine-tuning on our** [**80GB A100 Colab notebook**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt_oss_(20B)_500K_Context_Fine_tuning.ipynb) > **.** We’ve significantly improved how Unsloth handles memory usage patterns, speed, and context lengths: * **60% lower VRAM use** with **3.2x longer context** via Unsloth’s new [fused and chunked cross-entropy](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-loss-refactoring-chunk-and-fuse) loss, with no degradation in speed or accuracy * Enhanced activation offloading in Unsloth’s [**Gradient Checkpointing**](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-gradient-checkpointing-enhanced) * Collabing with Stas Bekman from Snowflake on [Tiled MLP](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#tiled-mlp-unlocking-500k) , enabling 2× more contexts Unsloth’s algorithms allows gpt-oss-20b QLoRA (4bit) with 290K context possible on a H100 with no accuracy loss, and 500K+ with Tiled MLP enabled, altogether delivering >**6.4x longer context lengths.** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8Ha930qR5XXBOK7M7oiy%252Fline_chart_light_tiled.png%3Falt%3Dmedia%26token%3D51467f68-a77b-4037-b9d9-e668223868c5&width=768&dpr=3&quality=100&sign=74316519&sv=2) ### [](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-loss-refactoring-chunk-and-fuse) 📐 Unsloth Loss Refactoring: Chunk & Fuse Our new fused loss implementation adds **dynamic sequence chunking**: instead of computing language model head logits and cross-entropies over the entire sequence at once, we process manageable slices along the flattened sequence dimension. This cuts peak memory from GBs to a smaller chunk sizes. Each chunk still runs a fully fused forward + backward pass via `torch.func.grad_and_value` , and retains mixed precision accuracy by upcasting to float32 if necessary. **These changes do not degrade training speed or accuracy.** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FFF43WA1X8Y4vADBrCi8T%252Fline_chart_light.png%3Falt%3Dmedia%26token%3D7afc7f73-bc54-403a-9674-8a16841ec659&width=768&dpr=3&quality=100&sign=43904d88&sv=2) The key innovation is that the **chunk size is chosen automatically at runtime** based on available VRAM. * If you have more free VRAM, larger chunks are used for faster runs * If you have less VRAM, it increases the number of chunks to avoid memory blowouts. This **removes manual tuning** and keeps our algorithm robust across old and new GPUs, workloads and different sequence lengths. Due to automatic tuning, **smaller contexts will use more VRAM** (fewer chunks) to **avoid unnecessary overhead**. For the plots above, we adjust the number of loss chunks to reflect realistic VRAM tiers. With 80GB VRAM, this yields >3.2× longer contexts. ### [](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-gradient-checkpointing-enhancements) 🏁 Unsloth Gradient Checkpointing Enhancements Our [Unsloth Gradient Checkpointing](https://unsloth.ai/blog/long-context) algorithm, **introduced in April 2024**, quickly became popular and the standard across the industry, having been integrated into most training packages nowadays. It offloads activations to CPU RAM which allowed 10x longer context lengths. Our new enhancements uses CUDA Streams and other tricks to add at most **0.1%** training overhead with no impact on accuracy. Previously it added 1 to 3% training overhead. By offloading activations as soon as they are produced, we minimize peak activation footprint and free GPU memory exactly when it’s needed. This sharply reduces memory pressure in long-context or large-batch training, where a single decoder layer’s activations can exceed 2 GB. > **Thus, Unsloth’s new algorithms & Gradient Checkpointing contributes to most improvements (3.2x), enabling 290k-context-length QLoRA GPT-OSS fine-tuning on a single H100.** ### [](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#tiled-mlp-unlocking-500k) 🔓 Tiled MLP: Unlocking 500K+ With help from [Stas Bekman](https://x.com/StasBekman) (Snowflake), we integrated Tiled MLP from Snowflake’s Arctic Long Sequence Training [paper](https://arxiv.org/abs/2506.13996) and blog post. TiledMLP reduces activation memory and enables much longer sequence lengths by tiling hidden states along the sequence dimension before heavy MLP projections. **We also introduce a few quality-of-life improvements:** We preserve RNG state across tiled forward recomputations so dropout and other stochastic ops are consistent between forward and backward replays. This keeps nested checkpointed computations stable and numerically identical. Our implementation auto patches any module named or typed as `mlp`, so **nearly all models with MLP modules are supported out of the box for Tiled MLP.** **Tradeoffs to keep in mind** TiledMLP saves VRAM at the cost of extra forward passes. Because it lives inside a checkpointed transformer block and is itself written in a checkpoint style, it effectively becomes a nested checkpoint: one **MLP now performs ~3 forward passes and 1 backward pass per step**. In return, we can drop almost all intermediate MLP activations from VRAM while still supporting extremely long sequences. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FdeOJEEqucGYtbXbb7nqB%252Fbaseline_vs_unsloth_spike.png%3Falt%3Dmedia%26token%3D3b1cdfd3-dd24-4c94-b7ec-5d1366464afb&width=768&dpr=3&quality=100&sign=29b337f3&sv=2) The plots compare active memory timelines for a single decoder layer’s forward and backward during a long-context training step, without Tiled MLP (left) and with it (right). Without Tiled MLP, peak VRAM occurs during the MLP backward; with Tiled MLP, it shifts to the fused loss calculation. We see ~40% lower VRAM usage, and because the fused loss auto chunks dynamically based on available VRAM, the peak with Tiled MLP would be even smaller on smaller GPUs. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FUCx0X7S5FvaD3hUsma5j%252Fbaseline_vs_unsloth_nospike.png%3Falt%3Dmedia%26token%3Da81b8639-21d0-43aa-a837-8209949e8742&width=768&dpr=3&quality=100&sign=78c626c&sv=2) To show cross-entropy loss is not the new bottleneck, we fix its chunk size instead of choosing it dynamically and then double the number of chunks. This significantly reduces the loss-related memory spikes. The max memory now occurs during backward in both cases, and overall timing is similar, though Tiled MLP adds a small overhead: one large GEMM becomes many sequential matmuls, plus the extra forward pass mentioned above. Overall, the trade-off is worth it: without Tiled MLP, long-context training can require roughly 2× the memory usage, while with **Tiled MLP a single GPU pays only about a 1.3× increase in step time for the same context length.** **Enabling Tiled MLP in Unsloth:** Just set `unsloth_tiled_mlp = True` in `from_pretrained` and Tiled MLP is enabled. We follow the same logic as the Arctic paper and choose `num_shards = ceil(seq_len/hidden_size)`. Each tile will operate on sequence lengths which are the same size of the hidden dimension of the model to balance throughput and memory savings. We also discussed how Tiled MLP effectively does 3 forward passes and 1 backward, compared to normal gradient checkpointing which does 2 forward passes and 1 backward with Stas Bekman and [DeepSpeed](https://github.com/deepspeedai/DeepSpeed/pull/7664) provided a doc update for Tiled MLP within DeepSpeed. Next time fine-tuning runs out of memory, try turning on `unsloth_tiled_mlp = True`. This should save some VRAM as long as the context length is longer than the LLM's hidden dimension. * * * **With our latest update, it is possible to now reach 1M context length with a smaller model on a single GPU!** **Try 500K-context gpt-oss-20b fine-tuning on our** [**80GB A100 Colab notebook**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt_oss_(20B)_500K_Context_Fine_tuning.ipynb) **.** If you've made it this far, we're releasing a new blog on our latest improvements in training speed this week so stay tuned by joining our [Reddit r/unsloth](https://www.reddit.com/r/unsloth/) or our Docs. [PreviousNew 3x Faster Training](https://unsloth.ai/docs/blog/3x-faster-training-packing) [NextQuantization-Aware Training](https://unsloth.ai/docs/blog/quantization-aware-training-qat) Last updated 7 months ago Was this helpful? * [📐 Unsloth Loss Refactoring: Chunk & Fuse](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-loss-refactoring-chunk-and-fuse) * [🏁 Unsloth Gradient Checkpointing Enhancements](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#unsloth-gradient-checkpointing-enhancements) * [🔓 Tiled MLP: Unlocking 500K+](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning#tiled-mlp-unlocking-500k) Was this helpful? Copy # Original Unsloth version released April 2024 - LGPLv3 Licensed class Unsloth_Offloaded_Gradient_Checkpointer(torch.autograd.Function): @staticmethod @torch_amp_custom_fwd def forward(ctx, forward_function, hidden_states, *args): ctx.device = hidden_states.device saved_hidden_states = hidden_states.to("cpu", non_blocking = True) with torch.no_grad(): output = forward_function(hidden_states, *args) ctx.save_for_backward(saved_hidden_states) ctx.forward_function, ctx.args = forward_function, args return output @staticmethod @torch_amp_custom_bwd def backward(ctx, dY): (hidden_states,) = ctx.saved_tensors hidden_states = hidden_states.to(ctx.device, non_blocking = True).detach() hidden_states.requires_grad_(True) with torch.enable_grad(): (output,) = ctx.forward_function(hidden_states, *ctx.args) torch.autograd.backward(output, dY) return (None, hidden_states.grad,) + (None,)*len(ctx.args) Show all 23 lines Copy model, tokenizer = FastLanguageModel.from_pretrained( ..., unsloth_tiled_mlp = True, ) --- # How to Run Diffusion Image GGUFs in ComfyUI | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/comfyui.md) . ComfyUI is an open-source diffusion model GUI, API, and backend that uses a node-based (graph/flowchart) interface. [ComfyUI](https://github.com/comfyanonymous/ComfyUI) is the most popular way to run workflows for image models like Qwen-Image-Edit or FLUX. GGUF is of the best and efficient formats for running diffusion models locally, and [Unsloth Dynamic](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs) GGUFs uses smart quantization to preserve accuracy even at low-bits. You'll learn how to install ComfyUI (Windows, Linux, macOS), build workflows, and tune [hyperparameters](https://unsloth.ai/docs/blog/comfyui#workflow-and-hyperparameters-1) in this step-by-step tutorial. #### [](https://unsloth.ai/docs/blog/comfyui#prerequisites-and-requirements) Prerequisites & Requirements You don’t need a GPU to run diffusion GGUFs, just a CPU with RAM. VRAM isn’t required but will make inference much faster. For best results, ensure your total usable memory (RAM + VRAM / unified) is slightly larger than the GGUF size; for example, the 4-bit (Q4\_K\_M) `unsloth/Qwen-Image-Edit-2511-GGUF` is 13.1 GB, so you should have at least ~13.2 GB of combined memory. You can find all Unsloth diffusion GGUFs in [our Collection](https://huggingface.co/collections/unsloth/unsloth-diffusion-ggufs) . We recommend at least 3-bit quantization for diffusion models, since their layers, especially the vision components, are very sensitive to quantization. Unsloth Dynamic quants upcasts important layers to recover as much accuracy as possible. [](https://unsloth.ai/docs/blog/comfyui#comfyui-tutorial) 📖 ComfyUI Tutorial ---------------------------------------------------------------------------------- ComfyUI represents the entire image generation pipeline as a graph of connected nodes. This guide will focus on machines with CUDA, but instructions to build with on Apple or CPU are similar. ### [](https://unsloth.ai/docs/blog/comfyui#id-1.-install-and-setup) #1. Install & Setup To install ComfyUI, you can download the desktop app on Windows or Mac devices [here](https://www.comfy.org/download) . Otherwise, to setup ComfyUI for running GGUF models run the following: Copy mkdir comfy_ggufs cd comfy_ggufs python -m venv .venv source .venv/bin/activate git clone https://github.com/comfyanonymous/ComfyUI.git cd ComfyUI pip install -r requirements.txt cd custom_nodes git clone https://github.com/city96/ComfyUI-GGUF cd ComfyUI-GGUF pip install -r requirements.txt cd ../.. ### [](https://unsloth.ai/docs/blog/comfyui#id-2.-download-models) #2. Download Models Diffusion models typically need 3 models. A Variational AutoEncoder (VAE) that encodes image pixel space to latent space, a text encoder to translate text to input embeddings, and the actual diffusion transformer. You can find all Unsloth diffusion GGUFs in our [Collection here](https://huggingface.co/collections/unsloth/unsloth-diffusion-ggufs) . Both the diffusion model and text encoder can be GGUF format while we typically use safetensors for the vae. Let's download the models we will use. See GGUF uploads for: [Qwen-Image-Edit-2511](https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF) , [FLUX.2-dev](https://huggingface.co/unsloth/FLUX.2-dev-GGUF) and [Qwen-Image-Layered](https://huggingface.co/unsloth/Qwen-Image-Layered-GGUF) The format of the vae and diffusion model might be different than the diffusers checkpoints. Only use checkpoints that are compatible with ComfyUI. These files must be in the correct folders for ComfyUI to see them. In addition the vision tower in the mmproj file must use the same prefix as the text encoder. Download reference images to be used later as well. #### [](https://unsloth.ai/docs/blog/comfyui#workflow-and-hyperparameters) Workflow and Hyperparameters You can also view our detailed [🎯 Workflow and Hyperparameters](https://unsloth.ai/docs/blog/comfyui#workflow-and-hyperparameters-1) Guide. Navigate to the ComfyUI directory and run: This will launch a web server that allows you to access `https://127.0.0.1:8188` . If you are running this on the cloud, you'll need to make sure port forwarding is setup to access on your local machine. Workflows are saved as JSON files embedded in output images (PNG metadata) or as separate `.json` files. You can: * Drag & drop an image into ComfyUI to load its workflow * Export/import workflows via the menu * Share workflows as JSON files Below are two examples of FLUX 2 json files which you can download and use: 15KB [unsloth\_flux2\_t2i\_gguf.json](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FpDft2nBJR3D1ti1zxr9v%2Funsloth_flux2_t2i_gguf.json?alt=media&token=43f65886-0a81-4bad-b6dd-f6d4daa89a9b) Download[Open](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FpDft2nBJR3D1ti1zxr9v%2Funsloth_flux2_t2i_gguf.json?alt=media&token=43f65886-0a81-4bad-b6dd-f6d4daa89a9b) 21KB [unsloth\_flux2\_i2i\_gguf.json](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FnCcWbaZDgETpESU0jrml%2Funsloth_flux2_i2i_gguf.json?alt=media&token=24c6926e-5b49-4aa5-93d7-9b6a63a0a3fa) Download[Open](https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FnCcWbaZDgETpESU0jrml%2Funsloth_flux2_i2i_gguf.json?alt=media&token=24c6926e-5b49-4aa5-93d7-9b6a63a0a3fa) Instead of setting up the workflow from scratch you can download the workflow here. Load it into the browser page by clicking the Comfy Logo -> File -> Open -> Then choose the `unsloth_flux2_t2i_gguf.json` file you just downloaded. It should look like the below: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FqoxBnRlnYrmzLfZshE1Z%252FScreenshot%2520from%25202025-12-29%252014-37-00.png%3Falt%3Dmedia%26token%3D1b1517b7-d44f-4e95-a5ed-759a4e0f74ec&width=768&dpr=3&quality=100&sign=d0ba0e31&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FWCVbmRbpijuj78M1rBtK%252FScreenshot%2520from%25202025-12-29%252021-41-52.png%3Falt%3Dmedia%26token%3D8378c519-1610-4752-bb71-4b95fdf00037&width=768&dpr=3&quality=100&sign=7f502f71&sv=2) This workflow is based on the official ComfyUI published workflow except it uses the GGUF loader extension, and is simplified to illustrate text to image functionality. ### [](https://unsloth.ai/docs/blog/comfyui#id-3.-inference) #3. Inference ComfyUI is highly customizable. You can mix models and create extremely complex pipelines. For a basic text to image setup we need to load the model, specify prompt and image details, and decide on a sampling strategy. **Upload Models + Set Prompt** We already downloaded the models, so we just need to pick the correct ones. For Unet Loader pick `flux2-dev-Q4_K_M.gguf`, for CLIPLoader pick `Mistral-Small-3.2-24B-Instruct-2506-UD-Q4_K_XL.gguf`, and for Load VAE pick `flux2-vae.safetensors`. You can set any prompt you'd like. Since classifier free guidance is baked into the model we do not need to specify a negative prompt. **Image Size + Sampler Parameters** Flux2-dev supports different image sizes. You can make rectangular shapes by setting the values of width and height. For sampler parameters, you can experiment with different samplers other than euler, and more or less sampling steps. Change the RandomNoise setting from randomize to fixed if you want to see how different settings change outputs. **Run** Click Run and an image will be generated in 45-60 seconds. That output image can be saved. The interesting part is that the metadata for the entire comfy workflow is saved in the image. You can share and anyone can see how it was created by loading it in the UI. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FfSKbmlievhyLzNmQxD88%252Funsloth_flux2_t2i_gguf.png%3Falt%3Dmedia%26token%3De9e1d8f0-777c-4083-823d-aeb3e77f5cf8&width=768&dpr=3&quality=100&sign=eab422bc&sv=2) **Multi Reference Generation** A key feature of Flux2 is multi reference generation where you can supply multiple images to use to help control generation. This time load the `unsloth_flux2_i2i_gguf.json`. We will use the same models, the only difference this time are extra nodes to select images to reference, which we've downloaded earlier. You'll notice the prompt refers to both `image 1` and `image 2` which are prompt anchors for the images. Once loaded click Run, and you'll see an output that creates our two unique sloth characters together while preserving their likeness. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FYyWL5368gfbwPAl8bhQ0%252Funsloth_flux2_i2i_gguf.png%3Falt%3Dmedia%26token%3D8857c028-5079-4eae-aec2-b02387bf2b23&width=768&dpr=3&quality=100&sign=c940351d&sv=2) [](https://unsloth.ai/docs/blog/comfyui#workflow-and-hyperparameters-1) 🎯 Workflow and Hyperparameters ------------------------------------------------------------------------------------------------------------ For text to image workflows we need to specify a prompt, sampling parameters, image size, guidance scale, and any optimization configs. #### [](https://unsloth.ai/docs/blog/comfyui#sampling) **Sampling** Sampling works differently from LLM's. Instead of sampling one token at a time we sample the whole image over multiple steps. Each step progressively "denoises" the image, which means that when you run for more steps, the image tends to be higher quality. There are also different sampling algorithms which range from first order and second order algorithms to deterministic and stochastic algorithms. For this tutorial we will use euler which a standard sampler that balances quality and speed. #### [](https://unsloth.ai/docs/blog/comfyui#guidance) **Guidance** Guidance is another important hyperparameter for diffusion models. There are many flavors of guidance but the two most widely used forms are **classifier free guidance (CFG)** and guidance distillation. The concept of classifier free guidance stems from [Classifier-Free Diffusion Guidance](https://arxiv.org/abs/2207.12598) . Historically you needed a separate classifier model to guide the model to match the input condition, but this paper actually shows CFG uses the difference between the model’s conditional and unconditional predictions to form a guidance direction. In practice it's not an unconditional prediction but a negative prompt prediction, meaning it's a prompt we definitely don't want and we should steer away from. When using CFG you do not need a separate model, but you need a second inference step from the unconditional or negative prompt. Other models have CFG baked in during training, but you can still set the strength of the guidance. This is separate from CFG since it does not need a second inference step, but it's still a tunable hyperparameter to set how strong its effect is. #### [](https://unsloth.ai/docs/blog/comfyui#conclusion) **Conclusion** Putting it all together, you set a prompt to tell the model what to produce, the text encoder encodes the text, the VAE encodes the image, both embeddings are stepped through the diffusion model according to the sampling parameters + guidance, and finally the output is decoded by the VAE which results in a usable image. ### [](https://unsloth.ai/docs/blog/comfyui#key-concepts-and-glossary) Key Concepts & Glossary * **Latent**: Compressed image representation (what the model operates on). * **Conditioning**: Text/image information that guides generation. * **Diffusion Model / UNet**: Neural network that performs the denoising. * **VAE**: Encoder/decoder between pixel space and latent space. * **CLIP (text encoder)**: Converts a prompt into embeddings. * **Sampler**: Algorithm that iteratively denoises the latent. * **Scheduler**: Controls the noise schedule across steps. * **Nodes**: Operations (load model, encode text, sample, decode, etc.). * **Edges**: Data flowing between nodes. Last updated 6 months ago Was this helpful? * [📖 ComfyUI Tutorial](https://unsloth.ai/docs/blog/comfyui#comfyui-tutorial) * [#1. Install & Setup](https://unsloth.ai/docs/blog/comfyui#id-1.-install-and-setup) * [#2. Download Models](https://unsloth.ai/docs/blog/comfyui#id-2.-download-models) * [#3. Inference](https://unsloth.ai/docs/blog/comfyui#id-3.-inference) * [🎯 Workflow and Hyperparameters](https://unsloth.ai/docs/blog/comfyui#workflow-and-hyperparameters-1) * [Key Concepts & Glossary](https://unsloth.ai/docs/blog/comfyui#key-concepts-and-glossary) Was this helpful? Copy cd models curl -L -C - -o vae/flux2-vae.safetensors \ https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors curl -L -C - -o text_encoders/Mistral-Small-3.2-24B-Instruct-2506-UD-Q4_K_XL.gguf \ https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF/resolve/main/Mistral-Small-3.2-24B-Instruct-2506-UD-Q4_K_XL.gguf curl -L -C - -o text_encoders/Mistral-Small-3.2-24B-Instruct-2506-mmproj-BF16.gguf \ https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF/resolve/main/mmproj-BF16.gguf curl -L -C - -o unet/flux2-dev-Q4_K_M.gguf \ https://huggingface.co/unsloth/FLUX.2-dev-GGUF/resolve/main/flux2-dev-Q4_K_M.gguf Copy curl -L -C - -o ../input/sloth1.jpg \ https://unsloth.ai/cgi/image/_1d5a5685-2d88-44ca-b50f-ba432cd646ef_9CGCY8lvw4D9JkOdueqsk.jpeg?width=1920&quality=80&format=auto curl -L -C - -o ../input/sloth2.jpg \ https://unsloth.ai/cgi/image/UnSloth_GPU_Front_-_Confetti_ArcSk-MR4MMN215UutOFZ.png?width=1920&quality=80&format=auto Copy python main.py --- # Fine-tuning Embedding Models with Unsloth Guide | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/embedding-finetuning.md) . Fine-tuning embedding models can largely improve retrieval and RAG performance on specific tasks. It aligns the model's vectors with your domain and the kind of 'similarity' that matters for your use case, which improves search, RAG, clustering, and recommendations on your data. Example: The headlines “Google launches Pixel 10” and “Qwen releases Qwen3” might be embedded as similar if you’re just labeling both as 'Tech,' but not similar if you’re doing semantic search because they’re about different things. Fine-tuning helps the model make the 'right' kind of similarity for your use case, reducing errors and improving results. [**Unsloth**](https://github.com/unslothai/unsloth) now supports training embedding, **classifier**, **BERT**, **reranker** models [**~1.8-3.3x faster**](https://unsloth.ai/docs/basics/embedding-finetuning#unsloth-benchmarks) with 20% less memory and 2x longer context than other Flash Attention 2 implementations - no accuracy degradation. EmbeddingGemma-300M works on just **3GB VRAM**. You can use your trained **model anywhere**: transformers, LangChain, Ollama, vLLM, llama.cpp etc. Unsloth uses [SentenceTransformers](https://github.com/huggingface/sentence-transformers) to support compatible models like Qwen3-Embedding, BERT and more. **Even if there's no notebook or upload, it’s still supported.** **We created free fine-tuning notebooks, with 3 main use-cases:** [EmbeddingGemma (300M)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/EmbeddingGemma_(300M).ipynb) [Qwen3-Embedding 4B](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_Embedding_(4B).ipynb) • [0.6B](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_Embedding_(0_6B).ipynb) [BGE M3](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/BGE_M3.ipynb) [ModernBERT](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert_classification.ipynb) - classification [All-MiniLM-L6-v2](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/All_MiniLM_L6_v2.ipynb) [ModernBERT-large](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert_classification.ipynb) * `All-MiniLM-L6-v2`: produce compact, domain-specific sentence embeddings for semantic search, retrieval, and clustering, tuned on your own data. * `tomaarsen/miriad-4.4M-split`: embed medical questions and biomedical papers for high-quality medical semantic search and RAG. * `electroglyph/technical`: better capture meaning and semantic similarity in technical text (docs, specs, and engineering discussions). You can view the rest of our uploaded models in [our collection here](https://huggingface.co/collections/unsloth/embedding-models) . > A huge thanks to Unsloth contributor [**electroglyph**](https://github.com/unslothai/unsloth/pull/3719) > , whose work was significant to support this. You can check out electroglyph’s custom models on Hugging Face [here](https://huggingface.co/electroglyph) > . ### [](https://unsloth.ai/docs/basics/embedding-finetuning#unsloth-features) 🦥 Unsloth Features * LoRA/QLoRA or full fine-tuning for embeddings, without needing to rewrite your pipeline * Best support for encoder-only `SentenceTransformer` models (with a `modules.json`) * Cross-encoder models are confirmed to train properly even under the fallback path * This release also supports `transformers v5` There is limited support for models without `modules.json` (we’ll auto-assign default `SentenceTransformers` pooling modules). If you’re doing something custom (custom heads, nonstandard pooling), double-check outputs like the pooled embedding behavior. Some models needed custom additions such as MPNet or DistilBERT were enabled by patching gradient checkpointing into the `transformers` models. ### [](https://unsloth.ai/docs/basics/embedding-finetuning#fine-tuning-workflow) 🛠️ Fine-tuning Workflow The new fine-tuning flow is centered around `FastSentenceTransformer`. Main save/push methods: * `save_pretrained()` Saves **LoRA adapters** to a local folder * `save_pretrained_merged()` Saves the **merged model** to a local folder * `push_to_hub()` Pushes **LoRA adapters** to Hugging Face * `push_to_hub_merged()` Pushes the **merged model** to Hugging Face **And one very important detail: Inference loading requires** `**for_inference=True**` `from_pretrained()` is similar to Lacker’s other fast classes, with **one exception**: * To load a model for **inference** using `FastSentenceTransformer`, you **must** pass: `for_inference=True` So your inference loads should look like: For Hugging Face authorization, if you run: inside the same virtualenv before calling the hub methods, then: * `push_to_hub()` and `push_to_hub_merged()` **don’t require a token argument**. ### [](https://unsloth.ai/docs/basics/embedding-finetuning#docs-internal-guid-c10bfa80-7fff-446e-714d-732eebcd72d6) ✅ Inference and Deploy Anywhere! Your fine-tuned Unsloth model can be used and deployed with all major tools: transformers, LangChain, Weaviate, sentence-transformers, Text Embeddings Inference (TEI), vLLM, and llama.cpp, custom embedding API, pgvector, FAISS/vector databases, and any RAG framework. There is no lock in as the fine-tuned model can later be downloaded locally on your own device. ### [](https://unsloth.ai/docs/basics/embedding-finetuning#unsloth-benchmarks) 📊 Unsloth Benchmarks Unsloth's advantages include speed for embedding fine-tuning! We show we are consistently **1.8 to 3.3x faster** on a wide variety of embedding models and on different sequence lengths from 128 to 2048 and longer. EmbeddingGemma-300M QLoRA works on just **3GB VRAM** and LoRA works on 6GB VRAM. Below are our Unsloth benchmarks in a heatmap vs. `SentenceTransformers` + Flash Attention 2 (FA2) for 4bit QLoRA. **For 4bit QLoRA, Unsloth is 1.8x to 2.6x faster:** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FQqagyYR6DebgX768A0HV%252Foutput%2816%29.png%3Falt%3Dmedia%26token%3De3ea6510-b129-401a-83ae-301d01865547&width=768&dpr=3&quality=100&sign=df622c31&sv=2) Below are our Unsloth benchmarks in a heatmap vs. `SentenceTransformers` + Flash Attention 2 (FA2) for 16bit LoRA. **For 16bit LoRA, Unsloth is 1.2x to 3.3x faster:** ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FTl12zuBg68ZPSyOC9hUe%252Foutput%2815%29.png%3Falt%3Dmedia%26token%3D47d7cade-7eac-4366-8011-7034de087431&width=768&dpr=3&quality=100&sign=2247dd61&sv=2) ### [](https://unsloth.ai/docs/basics/embedding-finetuning#model-support) 🔮 Model Support Here are some popular embedding models Unsloth supports (not all models are listed here): Most [common models](https://huggingface.co/models?library=sentence-transformers) are already supported. If there’s an encoder-only model you’d like that isn’t, feel free to open a [GitHub issue](https://github.com/unslothai/unsloth/issues) requesting it. [PreviousDistributed Data Parallel (DDP)](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp) [NextFaster MoE Training](https://unsloth.ai/docs/basics/faster-moe) Last updated 3 months ago Was this helpful? * [🦥 Unsloth Features](https://unsloth.ai/docs/basics/embedding-finetuning#unsloth-features) * [🛠️ Fine-tuning Workflow](https://unsloth.ai/docs/basics/embedding-finetuning#fine-tuning-workflow) * [✅ Inference and Deploy Anywhere!](https://unsloth.ai/docs/basics/embedding-finetuning#docs-internal-guid-c10bfa80-7fff-446e-714d-732eebcd72d6) * [📊 Unsloth Benchmarks](https://unsloth.ai/docs/basics/embedding-finetuning#unsloth-benchmarks) * [🔮 Model Support](https://unsloth.ai/docs/basics/embedding-finetuning#model-support) Was this helpful? Copy model = FastSentenceTransformer.from_pretrained( "sentence-transformers/all-MiniLM-L6-v2", for_inference=True, ) Copy hf auth login Copy # 1. Load a pretrained Sentence Transformer model model = SentenceTransformer(": Copy -v : Copy docker run -d -e JUPYTER_PORT=8000 \ -e JUPYTER_PASSWORD="mypassword" \ -e "SSH_KEY=$(cat ~/.ssh/container_key.pub)" \ -e USER_PASSWORD="unsloth2024" \ -p 8000:8000 -p 2222:22 \ -v $(pwd)/work:/workspace/work \ --gpus all \ unsloth/unsloth --- # Connect Anthropic to Unsloth: Run Claude Models in Local Chat | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/anthropic-claude.md) . Connect the Anthropic API to [Unsloth](https://github.com/unslothai/unsloth) to chat with Claude models including Claude Opus 4.7 directly alongside your local models in an open-source UI chat interface. This guide shows you how to create an Anthropic API key, add Anthropic as a provider in Unsloth, load Claude LLMs, and start chatting. Supported Claude models in Unsloth can also access advanced features such as thinking, [web search](https://unsloth.ai/docs/integrations/connections/anthropic-claude#web-search-and-thinking) , Anthropic [code execution](https://unsloth.ai/docs/integrations/connections/anthropic-claude#code-execution) , and [prompt caching](https://unsloth.ai/docs/integrations/connections/anthropic-claude#prompt-caching) to improve cost effiencey. ### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#setup) Setup 1 #### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#create-an-anthropic-api-key) Create an Anthropic API key Create an API key from the [Anthropic Console](https://console.anthropic.com/settings/keys) . Copy the key. You will paste it into Unsloth in the next step. 2 #### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#connect-anthropic-to-unsloth) Connect Anthropic to Unsloth ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FSzJ57mYBpBCavIKyksDy%252Fclaude_studio_api.gif%3Falt%3Dmedia%26token%3Dbd7b7da5-bb25-4f9d-bfb8-b63602c1f082&width=768&dpr=3&quality=100&sign=a2080164&sv=2) Next, connect Anthropic to Unsloth. 1. Open **Settings** → **Connections**, then click **Add Connection.** 2. Select the provider you want to add, then paste the API key you copied earlier. 3. Click **Reload Models** to refresh the list with models available to your account. 4. Choose the models you want to enable, then hit save. 3 #### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#ready-to-chat) Ready to Chat After saving the connection, select a Claude model under **Connected** in the model dropdown. Supported Claude models can expose extra controls including image generation, thinking, web search and code execution. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpOMKHdhunTepcJSvf8pA%252Fimage.png%3Falt%3Dmedia%26token%3Dca454678-a98b-4edb-b407-beeb879b4f76&width=768&dpr=3&quality=100&sign=3b641ca3&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#code-execution) Code Execution When enabled, supported Claude models can run code in Anthropic’s provider sandbox to solve problems, analyze data, and work with files. Claude uses Anthropic’s Code execution tool. Code execution appears in the response timeline as tool activity, alongside other tool calls. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F4W9Zq684jCgJy9kGvrkk%252FScreenshot%25202026-05-26%2520at%25206.10.12%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dc7b6a106-1a22-4120-a22d-f9ad8a9471c7&width=768&dpr=3&quality=100&sign=5befdb1c&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#prompt-caching) Prompt Caching Prompt caching reduces latency and cost when requests reuse the same long prefix. It is supported for compatible providers and servers, including Anthropic models. Use the **Prompt caching** setting in the side panel to control caching behavior for supported connections. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FKA6iU3qCFhUiq0KI16aU%252FScreenshot%25202026-05-26%2520at%25203.28.44%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D41a43794-792d-48f4-8774-8b9d85702dfc&width=768&dpr=3&quality=100&sign=3378c69e&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fq4csmTX5rS9isMkNkhTX%252FPrompt%2520Caching%2520Diagram%2520%281%29.png%3Falt%3Dmedia%26token%3Dade433bf-5eaf-4146-a266-525a85a6c98d&width=768&dpr=3&quality=100&sign=42ba571a&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#web-search-and-thinking) Web Search & Thinking Supported Claude models can use provider-side web search. The **Think** control appears when the selected model supports thinking. Depending on the model, this may expose different thinking levels or availability. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FC0Ed4kzN9h6c0ogEn6NT%252Fwebsearch%2520api.png%3Falt%3Dmedia%26token%3Dd3335222-1d9e-4021-9bf9-734c6acf1fc0&width=768&dpr=3&quality=100&sign=ecb70969&sv=2) The **Think** control adapts to the selected model: some models use an on/off toggle, while reasoning-effort models use model specific thinking levels. ### [](https://unsloth.ai/docs/integrations/connections/anthropic-claude#troubleshooting) Troubleshooting If Anthropic API fails to connect, check that the API key is valid and belongs to the correct Anthropic account. If a model does not appear after clicking **Load Models**, it may not be available for your account. You can enter the model ID manually or choose another model. [PreviousOpenAI](https://unsloth.ai/docs/integrations/connections/openai) [Nextllama.cpp / llama-server](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server) Last updated 1 month ago Was this helpful? * [Setup](https://unsloth.ai/docs/integrations/connections/anthropic-claude#setup) * [Code Execution](https://unsloth.ai/docs/integrations/connections/anthropic-claude#code-execution) * [Prompt Caching](https://unsloth.ai/docs/integrations/connections/anthropic-claude#prompt-caching) * [Web Search & Thinking](https://unsloth.ai/docs/integrations/connections/anthropic-claude#web-search-and-thinking) * [Troubleshooting](https://unsloth.ai/docs/integrations/connections/anthropic-claude#troubleshooting) Was this helpful? --- # Multi-GPU Fine-tuning with Distributed Data Parallel (DDP) | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp.md) . Let’s assume we have multiple GPUs, and we want to fine-tune a model using all of them! To do so, the most straightforward strategy is to use Distributed Data Parallel (DDP), which creates one copy of the model on each GPU device, feeding each copy distinct samples from the dataset during training and aggregating their contributions to weight updates per optimizer step. Why would we want to do this? Well, as we add more GPUs into the training process, we scale the number of samples our models train on per step, making each gradient update more stable and increasing our training throughput dramatically with each added GPU. Here’s a step-by-step guide on how to do this using Unsloth’s command-line interface (CLI)! **Note:** Unsloth DDP will work with any of your training scripts, not just via our CLI! More details below. #### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#install-unsloth-from-source) Install Unsloth from source We’ll clone Unsloth from GitHub and install it. Please consider using a [virtual environment](https://docs.python.org/3/tutorial/venv.html) ; we like to use `uv venv –python 3.12 && source .venv/bin/activate`, but any virtual environment creation tooling will do. Copy git clone https://github.com/unslothai/unsloth.git cd unsloth pip install . #### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#choose-target-model-and-dataset-for-finetuning) Choose target model and dataset for finetuning In this demo, we will fine-tune [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) on the [yahma/alpaca-cleaned](https://huggingface.co/datasets/yahma/alpaca-cleaned) chat dataset. This is a Supervised Fine-Tuning (SFT) workload that is commonly used when attempting to adapt a base model to a desired conversational style, or improve the model’s performance on a downstream task. ### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#use-the-unsloth-cli) Use the Unsloth CLI! First, let’s take a look at the help message built-in to the CLI (we’ve abbreviated here with “...” in various places for brevity): This should give you a sense of what options are available for you to pass into the CLI for training your model! For multi-GPU training (DDP in this case), we will use the [torchrun](https://docs.pytorch.org/docs/stable/elastic/run.html) launcher, which allows you to spin up multiple distributed training processes in single-node or multi-node settings. In our case, we will focus on the single-node (i.e., one machine) case with two H100 GPUs. Let’s also check our GPUs’ status by using the `nvidia-smi` command-line tool: Great! We have two H100 GPUs, as expected. Both are sitting at 0MiB memory usage as we’re currently not training anything, or have any model loaded into memory. To start your training run, issue a command like the following: If you have more GPUs, you may set `--nproc_per_node` accordingly to utilize them. **Note:** You can use the `torchrun` launcher with any of your Unsloth training scripts, including the [scripts](https://github.com/unslothai/notebooks/tree/main/python_scripts) converted from our free Colab notebooks, and DDP will be auto-enabled when training with >1 GPU! Taking a look again at `nvidia-smi` while training is in-flight: We can see that both GPUs are now using ~19GB of VRAM per H100 GPU! Inspecting the training logs, we see that we’re able to train at a rate of ~1.1 iterations/s. This training speed is ~constant even as we add more GPUs, so our training throughput increases ~linearly with the number of GPUs! ### [](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#training-metrics) Training metrics We ran a few short rank-16 LoRA fine-tunes on [unsloth/Llama-3.2-1B-Instruct](https://huggingface.co/unsloth/Llama-3.2-1B-Instruct) on the [yahma/alpaca-cleaned](https://huggingface.co/datasets/yahma/alpaca-cleaned) dataset to demonstrate the improved training throughput when using DDP training with multiple GPUs. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FdySJnhNUzVD3gsWmPqHR%252Funknown.png%3Falt%3Dmedia%26token%3D9905cccb-04c8-45b1-bfb1-680823713319&width=768&dpr=3&quality=100&sign=94c96e05&sv=2) The above figure compares training loss between two Llama-3.2-1B-Instruct LoRA fine-tunes over 500 training steps, with single GPU training (pink) vs. multi-GPU DDP training (blue). Notice that the loss curves match in scale and trend, but otherwise are a _bit_ different, since _the multi-GPU training processes twice as much training data per step_. This results in a slightly different training curve with less variability on a step-by-step basis. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fz4XgknzMgljaFInMEzHc%252Funknown.png%3Falt%3Dmedia%26token%3D4e28e2b1-8bc8-4049-983d-e4f980f3f4cf&width=768&dpr=3&quality=100&sign=caff1467&sv=2) The above figure plots training progress for the same two fine-tunes. Notice that the multi-GPU DDP training progresses through an epoch of the training data in half as many steps as single GPU training. This is because each GPU can process a distinct batch (of size `per_device_train_batch_size`) per step. However, the per-step timing for DDP training is slightly slower due to distributed communication for the model weight updates. As you increase the number of GPUs, the training throughput will continue to increase ~linearly (but with a small, but increasing penalty for the distributed comms). These same loss and training epoch progress behaviors hold for QLoRA fine-tunes, in which we loaded the base models in 4-bit precision in order to save additional GPU memory. This is particularly useful for training large models on limited amounts of GPU VRAM: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FUrCEgA7OBVhc8ICkMaP6%252Funknown.png%3Falt%3Dmedia%26token%3D0f5de3df-77df-4ee5-bf7a-68dead857c9a&width=768&dpr=3&quality=100&sign=52789357&sv=2) Training loss comparison between two Llama-3.2-1B-Instruct QLoRA fine-tunes over 500 training steps, with single GPU training (orange) vs. multi-GPU DDP training (purple). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8cG6rjmjeznNfgWrYdnG%252Funknown.png%3Falt%3Dmedia%26token%3Dd1c2c1fe-c117-49b5-8e9d-fdc01154cc01&width=768&dpr=3&quality=100&sign=9053ad9b&sv=2) Training progress comparison for the same two fine-tunes. [PreviousMulti-GPU Training](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth) [NextEmbedding Fine-tuning](https://unsloth.ai/docs/basics/embedding-finetuning) Last updated 7 months ago Was this helpful? * [Use the Unsloth CLI!](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#use-the-unsloth-cli) * [Training metrics](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp#training-metrics) Was this helpful? Copy $ python unsloth-cli.py --help usage: unsloth-cli.py [-h] [--model_name MODEL_NAME] [--max_seq_length MAX_SEQ_LENGTH] [--dtype DTYPE] [--load_in_4bit] [--dataset DATASET] [--r R] [--lora_alpha LORA_ALPHA] [--lora_dropout LORA_DROPOUT] [--bias BIAS] [--use_gradient_checkpointing USE_GRADIENT_CHECKPOINTING] … 🦥 Fine-tune your llm faster using unsloth! options: -h, --help show this help message and exit 🤖 Model Options: --model_name MODEL_NAME Model name to load --max_seq_length MAX_SEQ_LENGTH Maximum sequence length, default is 2048. We auto support RoPE Scaling internally! … 🧠 LoRA Options: These options are used to configure the LoRA model. --r R Rank for Lora model, default is 16. (common values: 8, 16, 32, 64, 128) --lora_alpha LORA_ALPHA LoRA alpha parameter, default is 16. (common values: 8, 16, 32, 64, 128) … 🎓 Training Options: --per_device_train_batch_size PER_DEVICE_TRAIN_BATCH_SIZE Batch size per device during training, default is 2. --per_device_eval_batch_size PER_DEVICE_EVAL_BATCH_SIZE Batch size per device during evaluation, default is 4. --gradient_accumulation_steps GRADIENT_ACCUMULATION_STEPS Number of gradient accumulation steps, default is 4. … Show all 36 lines Copy $ nvidia-smi Mon Nov 24 12:53:00 2025 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA H100 80GB HBM3 On | 00000000:04:00.0 Off | 0 | | N/A 32C P0 69W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ | 1 NVIDIA H100 80GB HBM3 On | 00000000:05:00.0 Off | 0 | | N/A 30C P0 68W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | No running processes found | +-----------------------------------------------------------------------------------------+ Show all 25 lines Copy # required: # --model_name # --dataset # optional; experiment with these: # --learning_rate, --max_seq_length, --per_device_train_batch_size, --gradient_accumulation_steps, --max_steps # to save the model at the end of training: # --save_model torchrun --nproc_per_node=2 unsloth-cli.py \ --model_name=Qwen/Qwen3-8B \ --dataset=yahma/alpaca-cleaned \ --learning_rate=2e-5 \ --max_seq_length=2048 \ --per_device_train_batch_size=1 \ --gradient_accumulation_steps=4 \ --max_steps=1000 \ --save_model Show all 17 lines Copy $ nvidia-smi Mon Nov 24 12:58:42 2025 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA H100 80GB HBM3 On | 00000000:04:00.0 Off | 0 | | N/A 38C P0 193W / 700W | 18903MiB / 81559MiB | 25% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ | 1 NVIDIA H100 80GB HBM3 On | 00000000:05:00.0 Off | 0 | | N/A 37C P0 199W / 700W | 18905MiB / 81559MiB | 28% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | 0 N/A N/A 4935 C ...und/unsloth/.venv/bin/python3 18256MiB | | 0 N/A N/A 4936 C ...und/unsloth/.venv/bin/python3 630MiB | | 1 N/A N/A 4935 C ...und/unsloth/.venv/bin/python3 630MiB | | 1 N/A N/A 4936 C ...und/unsloth/.venv/bin/python3 18258MiB | +-----------------------------------------------------------------------------------------+ Show all 28 lines --- # Fine-tuning LLMs with NVIDIA DGX Spark and Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth.md) . Unsloth enables local fine-tuning of LLMs with up to **200B parameters** on the NVIDIA DGX™ Spark. With 128 GB of unified memory, you can train massive models such as **gpt-oss-120b**, and run or deploy inference directly on DGX Spark. As shown at [OpenAI DevDay](https://x.com/UnslothAI/status/1976284209842118714) , gpt-oss-20b was trained with RL and Unsloth on DGX Spark to auto-win 2048. You can train using Unsloth in a Docker container or virtual environment on DGX Spark. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-ff5c4752dccb8f922b937f8e3b0db58e2d836507%252Funsloth%2520nvidia%2520dgx%2520spark.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=d11731a0&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-a8472482c49e1763378b609f8f537ca89df60260%252FNotebooks%2520on%2520dgx.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=288d9ddb&sv=2) In this tutorial, we’ll train gpt-oss-20b with RL using Unsloth notebooks after installing Unsloth on your DGX Spark. gpt-oss-120b will use around **68GB** of unified memory. After 1,000 steps and 4 hours of RL training, the gpt-oss model greatly outperforms the original on 2048, and longer training would further improve results. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-3bdcb0fda2ad188142e58f04c855b6dcfbd5ba94%252Fopenai%2520devday%2520unsloth%2520feature.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=cc268910&sv=2) You can watch Unsloth featured on OpenAI DevDay 2025 [here](https://youtu.be/1HL2YHRj270?si=8SR6EChF34B1g-5r&t=1080) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-4a8bd4ecc7ee3d123c19158df5dfdcec35df8532%252FScreenshot%25202025-10-13%2520at%25204.22.32%25E2%2580%25AFPM.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=417fc5f8&sv=2) gpt-oss trained with RL consistently outperforms on 2048. ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#step-by-step-tutorial) ⚡ Step-by-Step Tutorial 1 **Start with Unsloth Docker image for DGX Spark** First, build the Docker image using the DGX Spark Dockerfile which can be [found here](https://raw.githubusercontent.com/unslothai/notebooks/main/Dockerfile_DGX_Spark) . You can also run the below in a Terminal in the DGX Spark: Then, build the training Docker image using saved Dockerfile: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-7ebcf195c154b0e569115e1f9513cf002ee57b16%252Fdgx1.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=7f6afe05&sv=2) You can also click to see the full DGX Spark Dockerfile[](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#you-can-also-click-to-see-the-full-dgx-spark-dockerfile) 2 **Launch container** Launch the training container with GPU access and volume mounts: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-a67c36494f5c4ab4017748d490fb258655cd2378%252Fdgx2.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=14e268c5&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-b7758db087ab8b724049361781952b5ed154dfe8%252Fdgx5.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=6ee6379e&sv=2) 3 **Start Jupyter and Run Notebooks** Inside the container, start Jupyter and run the required notebook. You can use the Reinforcement Learning gpt-oss 20b to win 2048 [notebook here](https://github.com/unslothai/notebooks/blob/main/nb/gpt_oss_(20B)_Reinforcement_Learning_2048_Game_DGX_Spark.ipynb) . In fact all [Unsloth notebooks](https://docs.unsloth.ai/get-started/unsloth-notebooks) work in DGX Spark including the **120b** notebook! Just remove the installation cells. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-a8472482c49e1763378b609f8f537ca89df60260%252FNotebooks%2520on%2520dgx.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=288d9ddb&sv=2) The below commands can be used to run the RL notebook as well. After Jupyter Notebook is launched, open up the “`gpt_oss_20B_RL_2048_Game.ipynb`” ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-0862eed0acf0656ff0cb802b6aebc30892997e3b%252Fdgx6.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=d309acbd&sv=2) Don't forget Unsloth also allows you to [save and run](https://unsloth.ai/docs/basics/inference-and-deployment) your models after fine-tuning so you can locally deploy them directly on your DGX Spark after. Many thanks to [Lakshmi Ramesh](https://www.linkedin.com/in/rlakshmi24/) and [Barath Anandan](https://www.linkedin.com/in/barathsa/) from NVIDIA for helping Unsloth’s DGX Spark launch and building the Docker image. ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#unified-memory-usage) Unified Memory Usage gpt-oss-120b QLoRA 4-bit fine-tuning will use around **68GB** of unified memory. How your unified memory usage should look **before** (left) and **after** (right) training: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-e079a9aa8d853b319520fe0f0fbcca2e85b31ea6%252Fdgx7.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=bd11c1ff&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-c389a73a48ad059bbb92121b328fa7ccc61bee95%252Fdgx8.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=b3e1e4e9&sv=2) And that's it! Have fun training and running LLMs completely locally on your NVIDIA DGX Spark! ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#video-tutorials) Video Tutorials Thanks to Tim from [AnythingLLM](https://github.com/Mintplex-Labs/anything-llm) for providing a great fine-tuning tutorial with Unsloth on DGX Spark: [PreviousUnsloth Docker Guide](https://unsloth.ai/docs/blog/how-to-fine-tune-llms-with-unsloth-and-docker) [NextBlackwell, RTX 50 and Unsloth](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth) Last updated 7 months ago Was this helpful? * [⚡ Step-by-Step Tutorial](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#step-by-step-tutorial) * [Unified Memory Usage](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#unified-memory-usage) * [Video Tutorials](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth#video-tutorials) Was this helpful? Copy sudo apt update && sudo apt install -y wget wget -O Dockerfile "https://raw.githubusercontent.com/unslothai/notebooks/main/Dockerfile_DGX_Spark" Copy docker build -f Dockerfile -t unsloth-dgx-spark . Copy FROM nvcr.io/nvidia/pytorch:25.09-py3 # Set CUDA environment variables ENV CUDA_HOME=/usr/local/cuda-13.0/ ENV CUDA_PATH=$CUDA_HOME ENV PATH=$CUDA_HOME/bin:$PATH ENV LD_LIBRARY_PATH=$CUDA_HOME/lib64:$LD_LIBRARY_PATH ENV C_INCLUDE_PATH=$CUDA_HOME/include:$C_INCLUDE_PATH ENV CPLUS_INCLUDE_PATH=$CUDA_HOME/include:$CPLUS_INCLUDE_PATH # Install triton from source for latest blackwell support RUN git clone https://github.com/triton-lang/triton.git && \ cd triton && \ git checkout c5d671f91d90f40900027382f98b17a3e04045f6 && \ pip install -r python/requirements.txt && \ pip install . && \ cd .. # Install xformers from source for blackwell support RUN git clone --depth=1 https://github.com/facebookresearch/xformers --recursive && \ cd xformers && \ export TORCH_CUDA_ARCH_LIST="12.1" && \ python setup.py install && \ cd .. # Install unsloth and other dependencies RUN pip install unsloth unsloth_zoo bitsandbytes==0.48.0 transformers==4.56.2 trl==0.22.2 # Launch the shell CMD ["/bin/bash"] Copy docker run -it \ --gpus=all \ --net=host \ --ipc=host \ --ulimit memlock=-1 \ --ulimit stack=67108864 \ -v $(pwd):$(pwd) \ -v $HOME/.cache/huggingface:/root/.cache/huggingface \ -w $(pwd) \ unsloth-dgx-spark Copy NOTEBOOK_URL="https://raw.githubusercontent.com/unslothai/notebooks/refs/heads/main/nb/gpt_oss_(20B)_Reinforcement_Learning_2048_Game_DGX_Spark.ipynb" wget -O "gpt_oss_20B_RL_2048_Game.ipynb" "$NOTEBOOK_URL" jupyter notebook --ip=0.0.0.0 --port=8888 --no-browser --allow-root --- # Text-to-Speech (TTS) Fine-tuning Guide | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning.md) . Fine-tuning TTS models allows them to adapt to your specific dataset, use case, or desired style and tone. The goal is to customize these models to clone voices, adapt speaking styles and tones, support new languages, handle specific tasks and more. We also support **Speech-to-Text (STT)** models like OpenAI's Whisper. With [Unsloth](https://github.com/unslothai/unsloth) , you can fine-tune **any** TTS model (`transformers` compatible) 1.5x faster with 50% less memory than other implementations with Flash Attention 2. ⭐ **Unsloth supports any** `**transformers**` **compatible TTS model.** Even if we don’t have a notebook or upload for it yet, it’s still supported e.g., try fine-tuning Dia-TTS or Moshi. Zero-shot cloning captures tone but misses pacing and expression, often sounding robotic and unnatural. Fine-tuning delivers far more accurate and realistic voice replication. [Read more here](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-voice-models-vs.-zero-shot-voice-cloning) . ### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-notebooks) Fine-tuning Notebooks: We've also uploaded TTS models (original and quantized) to our [Hugging Face page](https://huggingface.co/collections/unsloth/text-to-speech-tts-models-68007ab12522e96be1e02155) . [Sesame-CSM (1B)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Sesame_CSM_(1B)-TTS.ipynb) [Orpheus-TTS (3B)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Orpheus_(3B)-TTS.ipynb) [Whisper Large V3](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Whisper.ipynb) (STT) [Spark-TTS (0.5B)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Spark_TTS_(0_5B).ipynb) [Llasa-TTS (1B)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llasa_TTS_(1B).ipynb) [Oute-TTS (1B)](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Oute_TTS_(1B).ipynb) If you notice that the output duration reaches a maximum of 10 seconds, increase`max_new_tokens = 125` from its default value of 125. Since 125 tokens corresponds to 10 seconds of audio, you'll need to set a higher value for longer outputs. ### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#choosing-and-loading-a-tts-model) Choosing and Loading a TTS Model For TTS, smaller models are often preferred due to lower latency and faster inference for end users. Fine-tuning a model under 3B parameters is often ideal, and our primary examples uses Sesame-CSM (1B) and Orpheus-TTS (3B), a Llama-based speech model. #### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#sesame-csm-1b-details) Sesame-CSM (1B) Details **CSM-1B** is a base model, while **Orpheus-ft** is fine-tuned on 8 professional voice actors, making voice consistency the key difference. CSM requires audio context for each speaker to perform well, whereas Orpheus-ft has this consistency built in. Fine-tuning from a base model like CSM generally needs more compute, while starting from a fine-tuned model like Orpheus-ft offers better results out of the box. To help with CSM, we’ve added new sampling options and an example showing how to use audio context for improved voice consistency. #### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#orpheus-tts-3b-details) Orpheus-TTS (3B) Details Orpheus is pre-trained on a large speech corpus and excels at generating realistic speech with built-in support for emotional cues like laughs and sighs. Its architecture makes it one of the easiest TTS models to utilize and train as it can be exported via llama.cpp meaning it has great compatibility across all inference engines. For unsupported models, you'll only be able to save the LoRA adapter safetensors. #### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#loading-the-models) Loading the models Because voice models are usually small in size, you can train the models using LoRA 16-bit or full fine-tuning FFT which may provide higher quality results. To load it in LoRA 16-bit: When this runs, Unsloth will download the model weights if you prefer 8-bit, you could use `load_in_8bit = True`, or for full fine-tuning set `full_finetuning = True` (ensure you have enough VRAM). You can also replace the model name with other TTS models. **Note:** Orpheus’s tokenizer already includes special tokens for audio output (more on this later). You do _not_ need a separate vocoder – Orpheus will output audio tokens directly, which can be decoded to a waveform. ### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#preparing-your-dataset) Preparing Your Dataset At minimum, a TTS fine-tuning dataset consists of **audio clips and their corresponding transcripts** (text). Let’s use the [_Elise_ dataset](https://huggingface.co/datasets/MrDragonFox/Elise) which is ~3 hour single-speaker English speech corpus. There are two variants: * [`MrDragonFox/Elise`](https://huggingface.co/datasets/MrDragonFox/Elise) – an augmented version with **emotion tags** (e.g. , ) embedded in the transcripts. These tags in angle brackets indicate expressions (laughter, sighs, etc.) and are treated as special tokens by Orpheus’s tokenizer * [`Jinsaryko/Elise`](https://huggingface.co/datasets/Jinsaryko/Elise) – base version with transcripts without special tags. The dataset is organized with one audio and transcript per entry. On Hugging Face, these datasets have fields such as `audio` (the waveform), `text` (the transcription), and some metadata (speaker name, pitch stats, etc.). We need to feed Unsloth a dataset of audio-text pairs. Instead of solely focusing on tone, cadence, and pitch, the priority should be ensuring your dataset is fully annotated and properly normalized. With some models like **Sesame-CSM-1B**, you might notice voice variation across generations using speaker ID 0 because it's a **base model**—it doesn’t have fixed voice identities. Speaker ID tokens mainly help maintain **consistency within a conversation**, not across separate generations. To get a consistent voice, provide **contextual examples**, like a few reference audio clips or prior utterances. This helps the model mimic the desired voice more reliably. Without this, variation is expected, even with the same speaker ID. **Option 1: Using Hugging Face Datasets library** – We can load the Elise dataset using Hugging Face’s `datasets` library: This will download the dataset (~328 MB for ~1.2k samples). Each item in `dataset` is a dictionary with at least: * `"audio"`: the audio clip (waveform array and metadata like sampling rate), and * `"text"`: the transcript string Orpheus supports tags like ``, ``, ``, ``, ``, ``, ``, ``, etc. For example: `"I missed you so much!"`. These tags are enclosed in angle brackets and will be treated as special tokens by the model (they match [Orpheus’s expected tags](https://github.com/canopyai/Orpheus-TTS) like `` and ``. During training, the model will learn to associate these tags with the corresponding audio patterns. The Elise dataset with tags already has many of these (e.g., 336 occurrences of “laughs”, 156 of “sighs”, etc. as listed in its card). If your dataset lacks such tags but you want to incorporate them, you can manually annotate the transcripts where the audio contains those expressions. **Option 2: Preparing a custom dataset** – If you have your own audio files and transcripts: * Organize audio clips (WAV/FLAC files) in a folder. * Create a CSV or TSV file with columns for file path and transcript. For example: * Use `load_dataset("csv", data_files="mydata.csv", split="train")` to load it. You might need to tell the dataset loader how to handle audio paths. An alternative is using the `datasets.Audio` feature to load audio data on the fly: Then `dataset[i]["audio"]` will contain the audio array. * **Ensure transcripts are normalized** (no unusual characters that the tokenizer might not know, except the emotion tags if used). Also ensure all audio have a consistent sampling rate (resample them if necessary to the target rate the model expects, e.g. 24kHz for Orpheus). In summary, for **dataset preparation**: * You need a **list of (audio, text)** pairs. * Use the HF `datasets` library to handle loading and optional preprocessing (like resampling). * Include any **special tags** in the text that you want the model to learn (ensure they are in `` format so the model treats them as distinct tokens). * (Optional) If multi-speaker, you could include a speaker ID token in the text or use a separate speaker embedding approach, but that’s beyond this basic guide (Elise is single-speaker). ### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-tts-with-unsloth) Fine-Tuning TTS with Unsloth Now, let’s start fine-tuning! We’ll illustrate using Python code (which you can run in a Jupyter notebook, Colab, etc.). **Step 1: Load the Model and Dataset** In all our TTS notebooks, we enable LoRA (16-bit) training and disable QLoRA (4-bit) training with: `load_in_4bit = False`. This is so the model can usually learn your dataset better and have higher accuracy. If memory is very limited or if dataset is large, you can stream or load in chunks. Here, 3h of audio easily fits in RAM. If using your own dataset CSV, load it similarly. **Step 2: Advanced - Preprocess the data for training (Optional)** We need to prepare inputs for the Trainer. For text-to-speech, one approach is to train the model in a causal manner: concatenate text and audio token IDs as the target sequence. However, since Orpheus is a decoder-only LLM that outputs audio, we can feed the text as input (context) and have the audio token ids as labels. In practice, Unsloth’s integration might do this automatically if the model’s config identifies it as text-to-speech. If not, we can do something like: The above is a simplification. In reality, to fine-tune Orpheus properly, you would need the _audio tokens as part of the training labels_. Orpheus’s pre-training likely involved converting audio to discrete tokens (via an audio codec) and training the model to predict those given the preceding text. For fine-tuning on new voice data, you would similarly need to obtain the audio tokens for each clip (using Orpheus’s audio codec). The Orpheus GitHub provides a script for data processing – it encodes audio into sequences of `` tokens. However, **Unsloth may abstract this away**: if the model is a FastModel with an associated processor that knows how to handle audio, it might automatically encode the audio in the dataset to tokens. If not, you’d have to manually encode each audio clip to token IDs (using Orpheus’s codebook). This is an advanced step beyond this guide, but keep in mind that simply using text tokens won’t teach the model the actual audio – it needs to match the audio patterns. Let's assume Unsloth provides a way to feed audio directly (for example, by setting `processor` and passing the audio array). If Unsloth does not yet support automatic audio tokenization, you might need to use the Orpheus repository’s `encode_audio` function to get token sequences for the audio, then use those as labels. (The dataset entries do have `phonemes` and some acoustic features which suggests a pipeline.) **Step 3: Set up training arguments and Trainer** We do 60 steps to speed things up, but you can set `num_train_epochs=1` for a full run, and turn off `max_steps=None`. Using a per\_device\_train\_batch\_size >1 may lead to errors if multi-GPU setup to avoid issues, ensure CUDA\_VISIBLE\_DEVICES is set to a single GPU (e.g., CUDA\_VISIBLE\_DEVICES=0). Adjust as needed. **Step 4: Begin fine-tuning** This will start the training loop. You should see logs of loss every 50 steps (as set by `logging_steps`). The training might take some time depending on GPU – for example, on a Colab T4 GPU, a few epochs on 3h of data may take 1-2 hours. Unsloth’s optimizations will make it faster than standard HF training. **Step 5: Save the fine-tuned model** After training completes (or if you stop it mid-way when you feel it’s sufficient), save the model. This ONLY saves the LoRA adapters, and not the full model. To save to 16bit or GGUF, scroll down! This saves the model weights (for LoRA, it might save only adapter weights if the base is not fully fine-tuned). If you used `--push_model` in CLI or `trainer.push_to_hub()`, you could upload it to Hugging Face Hub directly. Now you should have a fine-tuned TTS model in the directory. The next step is to test it out and if supported, you can use llama.cpp to convert it into a GGUF file. ### [](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-voice-models-vs.-zero-shot-voice-cloning) Fine-tuning Voice models vs. Zero-shot voice cloning People say you can clone a voice with just 30 seconds of audio using models like XTTS - no training required. That’s technically true, but it misses the point. Zero-shot voice cloning, which is also available in models like Orpheus and CSM, is an approximation. It captures the general **tone and timbre** of a speaker’s voice, but it doesn’t reproduce the full expressive range. You lose details like speaking speed, phrasing, vocal quirks, and the subtleties of prosody - things that give a voice its **personality and uniqueness**. If you just want a different voice and are fine with the same delivery patterns, zero-shot is usually good enough. But the speech will still follow the **model’s style**, not the speaker’s. For anything more personalized or expressive, you need training with methods like LoRA to truly capture how someone speaks. [PreviousFaster MoE Training](https://unsloth.ai/docs/basics/faster-moe) [NextDynamic 2.0 GGUFs](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs) Last updated 2 months ago Was this helpful? * [Fine-tuning Notebooks:](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-notebooks) * [Choosing and Loading a TTS Model](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#choosing-and-loading-a-tts-model) * [Preparing Your Dataset](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#preparing-your-dataset) * [Fine-Tuning TTS with Unsloth](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-tts-with-unsloth) * [Fine-tuning Voice models vs. Zero-shot voice cloning](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning#fine-tuning-voice-models-vs.-zero-shot-voice-cloning) Was this helpful? Copy from unsloth import FastModel model_name = "unsloth/orpheus-3b-0.1-pretrained" model, tokenizer = FastModel.from_pretrained( model_name, load_in_4bit=False # use 4-bit precision (QLoRA) ) Copy from datasets import load_dataset, Audio # Load the Elise dataset (e.g., the version with emotion tags) dataset = load_dataset("MrDragonFox/Elise", split="train") print(len(dataset), "samples") # ~1200 samples in Elise # Ensure all audio is at 24 kHz sampling rate (Orpheus’s expected rate) dataset = dataset.cast_column("audio", Audio(sampling_rate=24000)) Copy filename,text 0001.wav,Hello there! 0002.wav, I am very tired. Copy from datasets import Audio dataset = load_dataset("csv", data_files="mydata.csv", split="train") dataset = dataset.cast_column("filename", Audio(sampling_rate=24000)) Copy from unsloth import FastLanguageModel import torch dtype = None # None for auto detection. Float16 for Tesla T4, V100, Bfloat16 for Ampere+ load_in_4bit = False # Use 4bit quantization to reduce memory usage. Can be False. model, tokenizer = FastLanguageModel.from_pretrained( model_name = "unsloth/orpheus-3b-0.1-ft", max_seq_length= 2048, # Choose any for long context! dtype = dtype, load_in_4bit = load_in_4bit, #token = "hf_...", # use one if using gated models like meta-llama/Llama-2-7b-hf ) from datasets import load_dataset dataset = load_dataset("MrDragonFox/Elise", split = "train") Copy # Tokenize the text transcripts def preprocess_function(example): # Tokenize the text (keep the special tokens like intact) tokens = tokenizer(example["text"], return_tensors="pt") # Flatten to list of token IDs input_ids = tokens["input_ids"].squeeze(0) # The model will generate audio tokens after these text tokens. # For training, we can set labels equal to input_ids (so it learns to predict next token). # But that only covers text tokens predicting the next text token (which might be an audio token or end). # A more sophisticated approach: append a special token indicating start of audio, and let the model generate the rest. # For simplicity, use the same input as labels (the model will learn to output the sequence given itself). return {"input_ids": input_ids, "labels": input_ids} train_data = dataset.map(preprocess_function, remove_columns=dataset.column_names) Copy from transformers import TrainingArguments,Trainer,DataCollatorForSeq2Seq from unsloth import is_bfloat16_supported trainer = Trainer( model = model, train_dataset = dataset, args = TrainingArguments( per_device_train_batch_size = 1, gradient_accumulation_steps = 4, warmup_steps = 5, # num_train_epochs = 1, # Set this for 1 full training run. max_steps = 60, learning_rate = 2e-4, fp16 = not is_bfloat16_supported(), bf16 = is_bfloat16_supported(), logging_steps = 1, optim = "adamw_8bit", weight_decay = 0.01, lr_scheduler_type = "linear", seed = 3407, output_dir = "outputs", report_to = "none", # Use this for WandB etc ), ) Copy model.save_pretrained("lora_model") # Local saving tokenizer.save_pretrained("lora_model") # model.push_to_hub("your_name/lora_model", token = "...") # Online saving # tokenizer.push_to_hub("your_name/lora_model", token = "...") # Online saving --- # Connect llama.cpp to Unsloth: Run GGUFs with llama-server | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md) . Llama.cpp is an open-source inference engine for running GGUF models efficiently on local hardware, and [Unsloth](https://github.com/unslothai/unsloth) makes it easy to run those models directly into a open-source UI chat interface. By starting a local `llama-server`, you can serve a GGUF model from your machine or Hugging Face, connect it to Unsloth, and use it like any other external chat model. This guide walks through installing llama.cpp, launching `llama-server`, connecting it to Unsloth, enabling your model, and configuring prompt caching, context length, API keys, FA, and chat templates. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FhoLxJoFPUmjxslMPot7h%252Fexport-1779046578662.gif%3Falt%3Dmedia%26token%3D83b946dd-d8e1-4630-a143-b7b0e4d052a4&width=768&dpr=3&quality=100&sign=96afd0c7&sv=2) [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#setup) Setup ------------------------------------------------------------------------------------------------------------------------------ 1 ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#install-llama.cpp) Install llama.cpp Install llama.cpp first so you can run the `llama-server` command. Use one of the official install options: * Download a prebuilt [llama.cpp binary](https://github.com/ggml-org/llama.cpp/releases) * Build llama.cpp from [source](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md) After installing, check that llama-server works in your terminal: `llama-server --help` 2 ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#choose-a-gguf-model) Choose a GGUF model llama-server can load a local .gguf file or download a GGUF model from Hugging Face. To serve a Hugging Face GGUF repo directly, use the repo and quant name: `llama-server -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL` If you wish to load a local model, you can also follow the steps below. Start `llama-server` with the model you want to serve: This exposes an API endpoint at: `http://localhost:8080/v1` To require an API key, add: 3 ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#connect-llama.cpp-to-unsloth) Connect Llama.cpp to Unsloth Open **Settings → Connections**, then click **Add Connection**. Select **llama.cpp**, then enter your server details: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FU9TcWUfyBHEDWs91g8jx%252Fimage.png%3Falt%3Dmedia%26token%3D2a7a974a-9e2e-4bac-96f4-09e55e1767b9&width=768&dpr=3&quality=100&sign=1db694f1&sv=2) if you did not start llama-server with `--api-key`, leave the API key field empty. Enter the base URL of your server, e.g. `http://localhost:8080/v1` Click **Load Models** to fetch available model IDs, or enter model IDs manually if your server does not expose `/models`. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FXBMWShY6KWmxRJLwmieH%252Fimage.png%3Falt%3Dmedia%26token%3D1809ec26-84c6-4c61-b6a5-6dfa5bed769c&width=768&dpr=3&quality=100&sign=6ac7ee54&sv=2) Then, after you click **Add Connection,** The models you enabled will now appear under **Connected** in the **Select Model** dropdown. 4 ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#ready-to-chat) Ready to Chat After saving the connection, your llama.cpp model will appear under **Connection** in the model dropdown. Select it to start chatting through you **llama-server**. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FhoLxJoFPUmjxslMPot7h%252Fexport-1779046578662.gif%3Falt%3Dmedia%26token%3D83b946dd-d8e1-4630-a143-b7b0e4d052a4&width=768&dpr=3&quality=100&sign=96afd0c7&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#prompt-caching) Prompt Caching Prompt caching reduces latency and cost when requests reuse the same long prefix. Use the **Prompt caching** setting in the Unsloth side panel to control caching behaviour for supported connections. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fq4csmTX5rS9isMkNkhTX%252FPrompt%2520Caching%2520Diagram%2520%281%29.png%3Falt%3Dmedia%26token%3Dade433bf-5eaf-4146-a266-525a85a6c98d&width=768&dpr=3&quality=100&sign=42ba571a&sv=2) With llama.cpp, prompt caching is enabled by default and can be disabled when starting `llama-server` with: ### [](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#common-llama-server-arguments) **Common llama-server arguments** The example above only uses the required connection settings. You can add more llama-server arguments depending on your model and hardware. Common options include: For the full list of server arguments, see the official [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) . [PreviousAnthropic (Claude)](https://unsloth.ai/docs/integrations/connections/anthropic-claude) [NextvLLM](https://unsloth.ai/docs/integrations/connections/vllm) Last updated 1 month ago Was this helpful? * [Setup](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#setup) * [Install llama.cpp](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#install-llama.cpp) * [Choose a GGUF model](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#choose-a-gguf-model) * [Connect Llama.cpp to Unsloth](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#connect-llama.cpp-to-unsloth) * [Ready to Chat](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#ready-to-chat) * [Prompt Caching](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#prompt-caching) * [Common llama-server arguments](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#common-llama-server-arguments) Was this helpful? Copy llama-server \ --model /path/to/model.gguf \ --host 0.0.0.0 \ --port 8080 Copy --api-key 1234-myapi-key Copy --no-cache-prompt Copy --ctx-size 8192 \ # Set the context length --parallel 2 \ # Set the number of parallel slots --flash-attn on \ # Enable Flash Attention when supported --jinja \ # Use the model chat template --api-key 1234-key \ # Require an API key --no-cache-prompt # Disable prompt caching --- # Connect vLLM to Unsloth for Local Chat Inference | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/vllm.md) . Learn how to connect **vLLM to** [**Unsloth**](https://github.com/unslothai/unsloth) using vLLM’s **OpenAI-compatible API** so you can serve models and chat with them locally inside a open-source UI chat interface. This guide walks through installing vLLM, launching a local vLLM server, configuring the API base URL, loading available model IDs, and selecting your hosted vLLM model. By the end, your vLLM-served models will appear alongside local models, giving you a fast and flexible way to run external LLM inference from a UI chat interface. ### [](https://unsloth.ai/docs/integrations/connections/vllm#setup) Setup 1 #### [](https://unsloth.ai/docs/integrations/connections/vllm#install-vllm) Install vLLM Install vLLM first so you can run the `vllm serve` command. Follow the official [vLLM install guide](https://docs.vllm.ai/en/stable/getting_started/installation/) for your platform and hardware. After installing, check that vLLM works in your terminal: `vllm --help` 2 #### [](https://unsloth.ai/docs/integrations/connections/vllm#choose-a-model) Choose a model vLLM serves models from Hugging Face. For example, start a vLLM server with an Unsloth model: Copy vllm serve unsloth/gemma-4-26B-A4B-it \ --dtype auto This exposes an API endpoint at: `http://localhost:8000/v1` To require an API key, add: Copy --api-key token-abc123 3 #### [](https://unsloth.ai/docs/integrations/connections/vllm#connect-vllm-to-unsloth) Connect vLLM to Unsloth Open **Settings → Connections**, then click **Add Connection**. Select **vLLM**, then enter your server details. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FN80ObNBvpN7hDypheYkt%252Fimage.png%3Falt%3Dmedia%26token%3D481f4c1f-afb6-438d-ae35-85566f15e414&width=768&dpr=3&quality=100&sign=f6205043&sv=2) Enter your vLLM server details: * **API key:** leave empty unless you started vLLM with --api-key * **Base URL:** for example, http://localhost:8000/v1 * **Reasoning model:** enable this if the served model supports thinking * **Model IDs:** click **Load Models**, or enter custom IDs manually After you click **Add Connection**, the models you enabled will appear under **Connection** in the model dropdown. 4 #### [](https://unsloth.ai/docs/integrations/connections/vllm#ready-to-chat) Ready to Chat After saving the connection, your vLLM model will appear under **Connected** in the model dropdown. Select it to start chatting through your vLLM server. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpoXSmL3PcpprEy8CIT7S%252Fexport-1779046578662.gif%3Falt%3Dmedia%26token%3Dacf1de86-6fa4-466f-97f7-2d91662a124e&width=768&dpr=3&quality=100&sign=6e71dac1&sv=2) If your vLLM server is slow to respond (especially during model loading), you can adjust the timeout: Copy AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST=30 ### [](https://unsloth.ai/docs/integrations/connections/vllm#common-vllm-arguments) Common vLLM arguments The example above uses the core serving settings. You can add more vllm serve arguments depending on your model and hardware. Common options include: Copy vllm serve unsloth/gemma-4-26B-A4B-it \ --dtype auto \ --host 0.0.0.0 \ --port 8000 \ --api-key token-abc123 \ --max-model-len 8192 \ --gpu-memory-utilization 0.9 For the full list of vLLM server arguments, see the official vLLM [OpenAI-compatible server](https://docs.vllm.ai/en/stable/serving/openai_compatible_server/) docs. [Previousllama.cpp / llama-server](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server) [NextOllama](https://unsloth.ai/docs/integrations/connections/ollama) Last updated 1 month ago Was this helpful? * [Setup](https://unsloth.ai/docs/integrations/connections/vllm#setup) * [Common vLLM arguments](https://unsloth.ai/docs/integrations/connections/vllm#common-vllm-arguments) Was this helpful? --- # How to Run Local AI Models with OpenClaw | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/openclaw.md) . This guide will show how to use open LLMs locally with **OpenClaw by connecting it to Unsloth**. OpenClaw is an **open-source AI agent** interface that connects to a model to run tasks across your project. OpenClaw is able to works with any local model by connecting through **Unsloth’s OpenAI-compatible API**: including DeepSeek, Qwen, Gemma, and more. OpenClaw acts as the client, while Unsloth loads and serves models via a **local API**.After setup, OpenClaw will run against your local model through Unsloth, letting you use it directly as an **AI agent.** In this tutorial, we'll use [Qwen3.6](https://unsloth.ai/docs/models/qwen3.6) . [Connecting to OpenClaw](https://unsloth.ai/docs/integrations/openclaw#connecting-to-openclaw) [Quickstart](https://unsloth.ai/docs/integrations/openclaw#quickstart) In this tutorial, we’ll use `unsloth/Qwen3.6-27B-GGUF` in Unsloth and access it through OpenClaw. Prefer a different model? Swap in any other model by loading it in Unsloth and updating the configuration. ### [](https://unsloth.ai/docs/integrations/openclaw#installing-openclaw) Installing OpenClaw macOS, Linux, WSL Windows (PowerShell) Install OpenClaw using the official installer: `curl -fsSL https://openclaw.ai/install.sh | bash` This sets up OpenClaw and guides you through initial setup. Install OpenClaw using the official installer: `iwr -useb https://openclaw.ai/install.ps1 | iex` This sets up OpenClaw and guides you through initial setup. ### [](https://unsloth.ai/docs/integrations/openclaw#quickstart) ⚡ Quickstart After installing OpenClaw, we'll need to install Unsloth Studio to enable OpenClaw to serve and run inference of local models. 1. **Install or update** [**Unsloth Studio**](https://unsloth.ai/docs/new/studio) **.** Earlier versions don't expose the external API. See Installation. 2. **Launch Unsloth.** Note the port it starts on is usually `8000` or `8888`. You'll see it in the terminal output and in the browser URL (`http://localhost:PORT`). 3. **Load a model.** Click **New Chat**, pick or search a model (GGUF), and wait for it to finish loading. 4. **Connect OpenClaw.** Run `unsloth start openclaw` to launch OpenClaw with the loaded Unsloth model in a separate managed environment. Your normal OpenClaw configuration is left unchanged. ### [](https://unsloth.ai/docs/integrations/openclaw#launch-openclaw-with-unsloth-start) ⚙️ Launch OpenClaw with `unsloth start` OpenClaw can connect to a model already running in Unsloth Studio, or start one automatically when Unsloth is not running. Connect to a running Unsloth instance Once a model is loaded in Unsloth Studio, run: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FQm2hy6E5WvSzvmhQX0VL%252Fimage.png%3Falt%3Dmedia%26token%3De87ad9ac-8c43-4d20-9276-10eed404bd89&width=768&dpr=3&quality=100&sign=28270c1&sv=2) Unsloth launches OpenClaw in local TUI mode with the Unsloth provider, model, and context length configured inside a separate managed environment. Your normal OpenClaw setup is left unchanged. Once OpenClaw opens, give it a task such as: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fz93FYVhDJZKkBw3bY5X9%252Fimage.png%3Falt%3Dmedia%26token%3Df79db6e8-59ea-4ac7-8008-fe6e6124d977&width=768&dpr=3&quality=100&sign=62351f64&sv=2) #### [](https://unsloth.ai/docs/integrations/openclaw#start-a-model-automatically) Start a model automatically If Unsloth Studio is not already running, pass a model ID: Unsloth starts the model, launches OpenClaw, and stops the temporary server when you exit. To keep your OpenClaw state and return to the same session later, add `--persist` and choose a session key: Continue later using the same session key: See the complete [unsloth start](https://unsloth.ai/docs/integrations/unsloth-start) reference for named sessions, model loading, remote Unsloth servers, and advanced options. The rest of this guide covers the manual OpenClaw provider setup. ### [](https://unsloth.ai/docs/integrations/openclaw#manual-openclaw-provider-setup) Manual OpenClaw provider setup ### [](https://unsloth.ai/docs/integrations/openclaw#creating-an-api-key) 🔑 Creating an API key Keys are created from **Unsloth → Settings → API Keys**. 1. Open the sidebar, click your **Unsloth** avatar at the bottom-left. 2. Go to **Settings** → **API Keys**. 3. Enter a friendly name (e.g. `claude-code-macbook`). 4. _(Optional)_ Set an expiry. 5. Click **Create**. 6. **Copy the key immediately.** Unsloth stores only a hash and you won't be able to view it again. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FrWnqY5PgvZdbWK8pmlD5%252Fimage.png%3Falt%3Dmedia%26token%3D3507e7d7-b552-447f-b503-2418802b8f6e&width=768&dpr=3&quality=100&sign=d4395e2c&sv=2) All keys start with the `sk-unsloth-` prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with `401 Unauthorized`. ### [](https://unsloth.ai/docs/integrations/openclaw#connecting-to-openclaw) Connecting to OpenClaw OpenClaw reads its config from `~/.openclaw/openclaw.json`. Add (or merge) a `models` block with a `unsloth` provider pointing at Unsloth's Anthropic Messages API. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fs5g7Epfb2KjiQQf4MJf7%252Fopenclaw_unsloth_chat.png%3Falt%3Dmedia%26token%3Dc4de5f80-2a99-4527-bde0-c4e05e086f79&width=768&dpr=3&quality=100&sign=18f6ab91&sv=2) **Notes:** * `baseUrl` is the Unsloth origin with no path. OpenClaw talks to Unsloth over the Anthropic Messages API, and the Anthropic SDK appends `/v1/messages` itself, so do not add `/v1` here (a trailing `/v1` would send requests to `/v1/v1/messages`). * `api: "anthropic-messages"` tells OpenClaw to talk to Unsloth's `/v1/messages` endpoint. * `authHeader: true` sends your key as `Authorization: Bearer …`. * Set each model's `id` and `name` to the name you chose when loading the model in Unsloth. * If you're running Unsloth on a remote machine, replace `localhost:8888` with that machine's address (e.g. `http://10.0.0.42:8888`). ### [](https://unsloth.ai/docs/integrations/openclaw#optional-configure-model-behavior) Optional: configure model behavior OpenClaw connects through the model running in Unsloth. Runtime settings can be configured when starting the server. Use `--disable-tools` when driving OpenClaw (or any external coding agent). By default Unsloth Studio runs its own server-side tools, which swallows the agent's tool calls, so OpenClaw answers but never edits files. `--disable-tools` switches to passthrough, so OpenClaw's own tools are used. Use `--reasoning off` to turn thinking off, or `--reasoning on` to turn it on for models that support reasoning. This starts the server on `0.0.0.0:8888`, allowing other devices on your local network to connect. For more advanced runtime configuration, see the main [API tuning](https://unsloth.ai/docs/basics/api#unsloth-run-command) section. [PreviousHermes Agent](https://unsloth.ai/docs/integrations/hermes-agent) [NextOpenCode](https://unsloth.ai/docs/integrations/opencode) Last updated 4 days ago Was this helpful? * [Installing OpenClaw](https://unsloth.ai/docs/integrations/openclaw#installing-openclaw) * [⚡ Quickstart](https://unsloth.ai/docs/integrations/openclaw#quickstart) * [⚙️ Launch OpenClaw with unsloth start](https://unsloth.ai/docs/integrations/openclaw#launch-openclaw-with-unsloth-start) * [Manual OpenClaw provider setup](https://unsloth.ai/docs/integrations/openclaw#manual-openclaw-provider-setup) * [🔑 Creating an API key](https://unsloth.ai/docs/integrations/openclaw#creating-an-api-key) * [Connecting to OpenClaw](https://unsloth.ai/docs/integrations/openclaw#connecting-to-openclaw) * [Optional: configure model behavior](https://unsloth.ai/docs/integrations/openclaw#optional-configure-model-behavior) Was this helpful? Copy unsloth start openclaw Copy inspect this repo and summarize it in one sentence: https://github.com/unslothai/unsloth Copy unsloth start openclaw --model unsloth/Qwen3.6-27B-GGUF Copy unsloth start openclaw \ --model unsloth/Qwen3.6-27B-GGUF \ --persist \ agent --local \ --session-key my-session \ --message "Inspect this repository" Copy unsloth start openclaw \ --model unsloth/Qwen3.6-27B-GGUF \ --persist \ agent --local \ --session-key my-session \ --message "Continue" ~/.openclaw/openclaw.json Copy { "models": { "mode": "merge", "providers": { "unsloth": { "baseUrl": "http://localhost:8888", "apiKey": "sk-unsloth-xxxxxxxxxxxx", "api": "anthropic-messages", "models": [\ {\ "id": "unsloth/Qwen3.6-27B-GGUF",\ "name": "unsloth/Qwen3.6-27B-GGUF"\ }\ ], "authHeader": true } } } } Copy # Configure default generation behavior (--disable-tools passes OpenClaw's own tools through) unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ --disable-tools \ --reasoning off \ --temp 0.6 Copy # Allow connections from other devices unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ -H 0.0.0.0 \ -p 8888 --- # Vision Fine-tuning | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/vision-fine-tuning.md) . Fine-tuning vision models enables model to excel at certain tasks normal LLMs won't be as good as such as object/movement detection. **You can also train** [**VLMs with RL**](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl) **.** We have many free notebooks for vision fine-tuning: * [**Qwen3-VL**](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune) **(8B) Vision:** [**Notebook**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_VL_(8B)-Vision.ipynb) * [**Ministral 3**](https://unsloth.ai/docs/models/tutorials/ministral-3) : vision fine-tuning for general Q&A: [Notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Pixtral_(12B)-Vision.ipynb) One can concatenate general Q&A datasets with more niche datasets to make the finetune not forget base model skills. * **Gemma 3 (4B) Vision:** [Notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3_(4B)-Vision.ipynb) * **Llama 3.2 Vision** fine-tuning for radiography: [Notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(11B)-Vision.ipynb) How can we assist medical professionals in analyzing Xrays, CT Scans & ultrasounds faster. * **Qwen2.5 VL** fine-tuning for converting handwriting to LaTeX: [Notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5_VL_(7B)-Vision.ipynb) This allows complex math formulas to be easily transcribed as LaTeX without manually writing it. It is best to ensure your dataset has images of all the same size/dimensions. Use dimensions of 300-1000px to ensure your training does not take too long or use too many resources. ### [](https://unsloth.ai/docs/basics/vision-fine-tuning#disabling-vision-text-only-fine-tuning) Disabling Vision / Text-only fine-tuning To finetune vision models, we now allow you to select which parts of the mode to finetune. You can select to only finetune the vision layers, or the language layers, or the attention / MLP layers! We set them all on by default! Copy model = FastVisionModel.get_peft_model( model, finetune_vision_layers = True, # False if not finetuning vision layers finetune_language_layers = True, # False if not finetuning language layers finetune_attention_modules = True, # False if not finetuning attention layers finetune_mlp_modules = True, # False if not finetuning MLP layers r = 16, # The larger, the higher the accuracy, but might overfit lora_alpha = 16, # Recommended alpha == r at least lora_dropout = 0, bias = "none", random_state = 3407, use_rslora = False, # We support rank stabilized LoRA loftq_config = None, # And LoftQ target_modules = "all-linear", # Optional now! Can specify a list if needed modules_to_save=[\ "lm_head",\ "embed_tokens",\ ], ) ### [](https://unsloth.ai/docs/basics/vision-fine-tuning#vision-data-collator) Vision Data Collator We have a special data collator just for vision datasets: And the arguments for the data collator are: ### [](https://unsloth.ai/docs/basics/vision-fine-tuning#multi-image-training) Multi-image training In order to fine-tune or train models with multi-images the most straightforward change is to swap: with: Using map kicks in dataset standardization and arrow processing rules which can be strict and more complicated to define. ### [](https://unsloth.ai/docs/basics/vision-fine-tuning#dataset-for-vision-fine-tuning) Dataset for Vision Fine-tuning The dataset for fine-tuning a vision or multimodal model is similar to standard question & answer pair [datasets](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/datasets-guide) , but this time, they also includes image inputs. For example, the [Llama 3.2 Vision Notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(11B)-Vision.ipynb#scrollTo=vITh0KVJ10qX) uses a radiography case to show how AI can help medical professionals analyze X-rays, CT scans, and ultrasounds more efficiently. We'll be using a sampled version of the ROCO radiography dataset. You can access the dataset [here](https://www.google.com/url?q=https%3A%2F%2Fhuggingface.co%2Fdatasets%2Funsloth%2FRadiology_mini) . The dataset includes X-rays, CT scans and ultrasounds showcasing medical conditions and diseases. Each image has a caption written by experts describing it. The goal is to finetune a VLM to make it a useful analysis tool for medical professionals. Let's take a look at the dataset, and check what the 1st example shows: Image Caption ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-97d4489827403bd4795494f33d01a10979788c30%252Fxray.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=369d7ad8&sv=2) Panoramic radiography shows an osteolytic lesion in the right posterior maxilla with resorption of the floor of the maxillary sinus (arrows). To format the dataset, all vision finetuning tasks should be formatted as follows: We will craft an custom instruction asking the VLM to be an expert radiographer. Notice also instead of just 1 instruction, you can add multiple turns to make it a dynamic conversation. Let's convert the dataset into the "correct" format for finetuning: The first example is now structured like below: Before we do any finetuning, maybe the vision model already knows how to analyse the images? Let's check if this is the case! And the result: For more details, view our dataset section in the [notebook here](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(11B)-Vision.ipynb#scrollTo=vITh0KVJ10qX) . ### [](https://unsloth.ai/docs/basics/vision-fine-tuning#training-on-assistant-responses-only-for-vision-models-vlms) 🔎Training on assistant responses only for vision models, VLMs For language models, we can use `from unsloth.chat_templates import train_on_responses_only` as described previously. For vision models, use the extra arguments as part of `UnslothVisionDataCollator` just like before! See [Vision Data Collator](https://unsloth.ai/docs/basics/vision-fine-tuning#vision-data-collator) for more details on how to use the vision data collator. For example for Llama 3.2 Vision: [PreviousTool Calling Guide](https://unsloth.ai/docs/basics/tool-calling-guide-for-local-llms) [NextTroubleshooting & FAQs](https://unsloth.ai/docs/basics/troubleshooting-and-faqs) Last updated 2 months ago Was this helpful? * [Disabling Vision / Text-only fine-tuning](https://unsloth.ai/docs/basics/vision-fine-tuning#disabling-vision-text-only-fine-tuning) * [Vision Data Collator](https://unsloth.ai/docs/basics/vision-fine-tuning#vision-data-collator) * [Multi-image training](https://unsloth.ai/docs/basics/vision-fine-tuning#multi-image-training) * [Dataset for Vision Fine-tuning](https://unsloth.ai/docs/basics/vision-fine-tuning#dataset-for-vision-fine-tuning) * [🔎Training on assistant responses only for vision models, VLMs](https://unsloth.ai/docs/basics/vision-fine-tuning#training-on-assistant-responses-only-for-vision-models-vlms) Was this helpful? Copy from unsloth.trainer import UnslothVisionDataCollator from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( model = model, tokenizer = tokenizer, data_collator = UnslothVisionDataCollator(model, tokenizer), train_dataset = dataset, args = SFTConfig(...), ) Copy class UnslothVisionDataCollator: def __init__( self, model, processor, max_seq_length = None, # [Optional] We auto get this from `FastVisionModel.from_pretrained(max_seq_length = ...) formatting_func = None, # Function for transforming the text resize = "min", # Can be (10, 10) or "min" to resize to fit the model's default image_size or "max" # for no resizing and leave image intact ignore_index = -100, # [Optional] Default is -100 # from unsloth.chat_templates import train_on_responses_only # trainer = train_on_responses_only( # trainer, # instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n", # response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n", # ) train_on_responses_only = False, # EQUIVALENT to train_on_responses_only for LLMs instruction_part = None, # EQUIVALENT to train_on_responses_only(instruction_part = ...) response_part = None, # EQUIVALENT to train_on_responses_only(response_part = ...) force_match = True, # Match newlines as well! num_proc = None, # [Optional] WIll auto select number of GPUs completion_only_loss = True, # [Optional] Ignores padding vision tokens - should always be True! pad_to_multiple_of = None, # [Optional] For data collator padding resize_dimension = 0, # can be 0, 1, 'max' or 'min' # (max resizes based on the max of height width, min the min size, 0 the first dim, etc) snap_to_patch_size = False, # [Optional] Force image to be a multiple of the patch size ) Show all 29 lines Copy ds_converted = ds.map( convert_to_conversation, ) Copy ds_converted = [convert_to_conversation(sample) for sample in dataset] Copy Dataset({ features: ['image', 'image_id', 'caption', 'cui'], num_rows: 1978 }) Copy [\ { "role": "user",\ "content": [{"type": "text", "text": instruction}, {"type": "image", "image": image} ]\ },\ { "role": "assistant",\ "content": [{"type": "text", "text": answer} ]\ },\ ] Copy instruction = "You are an expert radiographer. Describe accurately what you see in this image." def convert_to_conversation(sample): conversation = [\ { "role": "user",\ "content" : [\ {"type" : "text", "text" : instruction},\ {"type" : "image", "image" : sample["image"]} ]\ },\ { "role" : "assistant",\ "content" : [\ {"type" : "text", "text" : sample["caption"]} ]\ },\ ] return { "messages" : conversation } pass Show all 16 lines Copy converted_dataset = [convert_to_conversation(sample) for sample in dataset] Copy converted_dataset[0] Copy {'messages': [{'role': 'user',\ 'content': [{'type': 'text',\ 'text': 'You are an expert radiographer. Describe accurately what you see in this image.'},\ {'type': 'image',\ 'image': }]},\ {'role': 'assistant',\ 'content': [{'type': 'text',\ 'text': 'Panoramic radiography shows an osteolytic lesion in the right posterior maxilla with resorption of the floor of the maxillary sinus (arrows).'}]}]} Copy FastVisionModel.for_inference(model) # Enable for inference! image = dataset[0]["image"] instruction = "You are an expert radiographer. Describe accurately what you see in this image." messages = [\ {"role": "user", "content": [\ {"type": "image"},\ {"type": "text", "text": instruction}\ ]}\ ] input_text = tokenizer.apply_chat_template(messages, add_generation_prompt = True) inputs = tokenizer( image, input_text, add_special_tokens = False, return_tensors = "pt", ).to("cuda") from transformers import TextStreamer text_streamer = TextStreamer(tokenizer, skip_prompt = True) _ = model.generate(**inputs, streamer = text_streamer, max_new_tokens = 128, use_cache = True, temperature = 1.5, min_p = 0.1) Show all 23 lines Copy This radiograph appears to be a panoramic view of the upper and lower dentition, specifically an Orthopantomogram (OPG). * The panoramic radiograph demonstrates normal dental structures. * There is an abnormal area on the upper right, represented by an area of radiolucent bone, corresponding to the antrum. **Key Observations** * The bone between the left upper teeth is relatively radiopaque. * There are two large arrows above the image, suggesting the need for a closer examination of this area. One of the arrows is in a left-sided position, and the other is in the right-sided position. However, only Copy class UnslothVisionDataCollator: def __init__( self, ... # from unsloth.chat_templates import train_on_responses_only # trainer = train_on_responses_only( # trainer, # instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n", # response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n", # ) train_on_responses_only = False, # EQUIVALENT to train_on_responses_only for LLMs instruction_part = None, # EQUIVALENT to train_on_responses_only(instruction_part = ...) response_part = None, # EQUIVALENT to train_on_responses_only(response_part = ...) force_match = True, # Match newlines as well! ) Copy UnslothVisionDataCollator( model, tokenizer, ... train_on_responses_only = True, instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n", response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n", ... ) --- # Fine-Tuning LLMs on NVIDIA DGX Station with Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/dgx-station.md) . You can now train LLMs locally on your NVIDIA DGX Station with [Unsloth](https://github.com/unslothai/unsloth) . DGX Station has more than **~200GB VRAM** and over **700GB of unified GPU / CPU memory** and combines a Grace CPU and a Blackwell GPU in a tightly connected system designed for large-scale AI workloads. Linked by NVLink-C2C, the CPU and GPU remain distinct but work together far more efficiently than in a traditional CPU-GPU setup. In this guide, we’ll use Unsloth notebooks train [Qwen3.5](https://unsloth.ai/docs/blog/dgx-station#qwen3.5-35b-a3b-fine-tuning) and [gpt-oss-120b](https://unsloth.ai/docs/blog/dgx-station#gpt-oss-120b-fine-tuning) on DGX Station. Thank you to NVIDIA for providing some early access DGX Station hardware to test Unsloth on! ### [](https://unsloth.ai/docs/blog/dgx-station#quickstart) Quickstart You will need `python3` installed and in particular the dev headers are needed. On our system we have `python 3.12` so we will install the 3.12 dev headers. Copy sudo apt update sudo apt install python3.12-dev Then create a fresh virtual environment to install [Unsloth](https://github.com/unslothai/unsloth) . This way we minimize dependency conflicts and preserve the state of the current working environment. Copy python3 -m venv .unsloth source .unsloth/bin/activate pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 First install `torch` from the `cuda 13` index otherwise we could get the CPU version or a mismatch in architecture and capabilities! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fw04Su0JZriUaQxD31wf0%252Funknown.png%3Falt%3Dmedia%26token%3D83e61cdb-74c3-42c4-a1ff-18cec3752c9e&width=768&dpr=3&quality=100&sign=73a9ac77&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F9bs6h6YxI2hqnqOz1bU0%252Funknown.png%3Falt%3Dmedia%26token%3De3e261b5-be18-4d49-9f38-526012add332&width=768&dpr=3&quality=100&sign=4bd7cb89&sv=2) Now we can install Unsloth: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FhQZznPQ8O9Wh3At6FclO%252Funknown.png%3Falt%3Dmedia%26token%3D34c8de6e-bef8-414c-8e1b-2913589c4b10&width=768&dpr=3&quality=100&sign=4f1bf758&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FdLZCFmln5LaUWtO6eC4A%252Funknown.png%3Falt%3Dmedia%26token%3Dce04e025-32c7-4847-ac35-bee1baf6259f&width=768&dpr=3&quality=100&sign=ccb6bc91&sv=2) Now lets install `xformers` and optionally build `flash-attention` from source. Both packages take time so please be patient while they build. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FnyczIn3YvXAPx5oIfZQQ%252Funknown.png%3Falt%3Dmedia%26token%3D1a2c5f7b-13c5-4f5e-b4c4-61df8d5fc653&width=768&dpr=3&quality=100&sign=7857435d&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FoupUFzx2pOG6l5B91Pw4%252Funknown.png%3Falt%3Dmedia%26token%3D009d2c73-5992-4593-8fd0-e7d813eda3ff&width=768&dpr=3&quality=100&sign=fb07e33f&sv=2) For Qwen 3.5 MoE we’ll want to download two kernel packages `flash-linear-attention` and `causal-conv1d` to make it fast. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F4xEY8k3jzxfOgMWAgJD7%252Funknown.png%3Falt%3Dmedia%26token%3D2b8bd62e-23cd-4bcf-a0af-6d161d1ec1a1&width=768&dpr=3&quality=100&sign=be483f7a&sv=2) If you don’t already have a notebook client, install one. For this guide we will use Jupyter Notebook: Finally we download the actual Unsloth notebooks to run. There are 250+ notebooks for LLM Training as well as Python scripts. ### [](https://unsloth.ai/docs/blog/dgx-station#training-tutorials) Training Tutorials Now we can launch Jupyter Notebook and navigate to the UI on a browser. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FP2seywdWvLHHQkdP8DGy%252Funknown.png%3Falt%3Dmedia%26token%3Dca1b5390-5eb8-416b-a3e9-d9df9b27fb0b&width=768&dpr=3&quality=100&sign=fd1facf4&sv=2) Copy and paste the `localhost` site with token parameter and paste into your browser. You should see something like: The `nb` folder has all the notebooks to run. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FSxN976oDM4WaG5EtpSc9%252Funknown.png%3Falt%3Dmedia%26token%3D7113ba12-5bcc-4bc6-9777-b9d4c440d0bf&width=768&dpr=3&quality=100&sign=181360ba&sv=2) #### [](https://unsloth.ai/docs/blog/dgx-station#qwen3.5-35b-a3b-training) Qwen3.5-35B-A3B Training Open the file `nb/Qwen3_5_MoE.ipynb`. Skip past the installation section since we already installed everything we need beforehand. Navigate to the Unsloth section and start executing cells from there. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fif8mAvc1au9Hl83IZzNm%252FDGX%2520Station.png%3Falt%3Dmedia%26token%3D1011c8a9-c6ba-48df-a726-d3bc3bc8e947&width=768&dpr=3&quality=100&sign=8809e641&sv=2) The notebook covers model setup, dataset preparation, and trainer configuration. Each step can take some time as we are downloading a very large model, initializing billions of weights, and further optimizing to make it run fast. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F4v0NmHdhiYCHFll8U8OD%252Funknown.png%3Falt%3Dmedia%26token%3D69e3d279-4d59-4439-802f-11bd02fe39d3&width=768&dpr=3&quality=100&sign=2d91fe9c&sv=2) Training is very fast with the default setting. On the DGX Station there is plenty of memory so you can play with the default training hyper parameters to really push the memory and compute. Once done training you can save the model for later, push the model to Hugging Face Hub to share with others, or export to a quantized format. #### [](https://unsloth.ai/docs/blog/dgx-station#gpt-oss-120b-training) gpt-oss-120b Training Open the file `nb/gpt-oss-(120B)_A100-Fine-tuning.ipynb`. Skip past the installation section since we already installed the prerequisites and navigate to the Unsloth section. We can start running the notebook from there. The notebook will use around 72 GB of GPU memory and take about 10 minutes. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8jYYievlemxDJBatevNV%252FDGX%2520Station%25202.png%3Falt%3Dmedia%26token%3Defef1a26-a170-4690-972f-1a7cde67e9ea&width=768&dpr=3&quality=100&sign=963c9c14&sv=2) Each cell can take some time to run as we need to download the model, initialize the weights, and further optimize for a fast experience. The notebook goes through dataset preprocessing and trainer setup. Once we get to the `trainer.train()` cell and execute training begins. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FOxuma3ZZeEbxZrAWIgnq%252FDGX%2520Station%25203.png%3Falt%3Dmedia%26token%3D17beb84e-eb56-4357-aee2-078c4db3eb84&width=768&dpr=3&quality=100&sign=d0a8d479&sv=2) Now that it’s complete we can save the model for later use, push to Hugging Face Hub to share with the world, or export it to GGUF format. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fy1UxtQ01avFK5BIkofwt%252Fimage.png%3Falt%3Dmedia%26token%3D8d137818-a3a6-4d00-a9fd-1e41ed0483a5&width=768&dpr=3&quality=100&sign=76c758e8&sv=2) Read more about NVIDIA's DGX Station at [https://www.nvidia.com/en-us/products/workstations/dgx-station/](https://www.nvidia.com/en-us/products/workstations/dgx-station/) [PreviousQuantization-Aware Training](https://unsloth.ai/docs/blog/quantization-aware-training-qat) [NextUnsloth Docker Guide](https://unsloth.ai/docs/blog/how-to-fine-tune-llms-with-unsloth-and-docker) Last updated 4 months ago Was this helpful? * [Quickstart](https://unsloth.ai/docs/blog/dgx-station#quickstart) * [Training Tutorials](https://unsloth.ai/docs/blog/dgx-station#training-tutorials) Was this helpful? Copy pip install unsloth Copy pip install --no-deps --no-build-isolation xformers==0.0.33.post1 # Optionally flash-attn # Clone and build (targets sm_100 for B300) git clone https://github.com/Dao-AILab/flash-attention cd flash-attention # B300 = sm_100, set arch explicitly TORCH_CUDA_ARCH_LIST="10.0" MAX_JOBS=8 pip install . --no-build-isolation cd .. Copy pip install --no-build-isolation flash-linear-attention causal_conv1d==1.6.0 Copy cd .. pip install notebook pip install ipywidgets Copy git clone https://github.com/unslothai/notebooks.git cd notebooks Copy jupyter notebook --- # Troubleshooting & FAQs | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/troubleshooting-and-faqs.md) . If you're still encountering any issues with versions or dependencies, please use our [Docker image](https://unsloth.ai/docs/get-started/install/docker) which will have everything pre-installed. **Try always to update Unsloth if you find any issues.** `pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth_zoo` ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#fine-tuning-a-new-model-not-supported-by-unsloth) Fine-tuning a new model not supported by Unsloth? Unsloth works with any model supported by `transformers`. If a model isn’t in our uploads or doesn’t run out of the box, it’s usually still supported, some newer models may just need a small manual tweak due to our optimizations. In most cases, you can enable compatibility by setting `trust_remote_code=True` in your fine-tuning script. Here’s an example using [DeepSeek-OCR](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune) : ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#running-in-unsloth-works-well-but-after-exporting-and-running-on-other-platforms-the-results-are-poo) Running in Unsloth works well, but after exporting & running on other platforms, the results are poor You might sometimes encounter an issue where your model runs and produces good results on Unsloth, but when you use it on another platform like Ollama or vLLM, the results are poor or you might get gibberish, endless/infinite generations _or_ repeated outputs**.** * The most common cause of this error is using an **incorrect chat template****.** It’s essential to use the SAME chat template that was used when training the model in Unsloth and later when you run it in another framework, such as llama.cpp or Ollama. When inferencing from a saved model, it's crucial to apply the correct template. * It might also be because your inference engine adds an unnecessary "start of sequence" token (or the lack of thereof on the contrary) so ensure you check both hypotheses! * **Use our conversational notebooks to force the chat template - this will fix most issues.** * Qwen-3 14B Conversational notebook [**Open in Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(14B)-Reasoning-Conversational.ipynb) * Gemma-3 4B Conversational notebook [**Open in Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3_(4B).ipynb) * Llama-3.2 3B Conversational notebook [**Open in Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(1B_and_3B)-Conversational.ipynb) * Phi-4 14B Conversational notebook [**Open in Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi_4-Conversational.ipynb) * Mistral v0.3 7B Conversational notebook [**Open in Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral_v0.3_(7B)-Conversational.ipynb) * **More notebooks in our** [**notebooks docs**](https://unsloth.ai/docs/get-started/unsloth-notebooks) ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#saving-to-gguf-vllm-16bit-crashes) Saving to GGUF / vLLM 16bit crashes You can try reducing the maximum GPU usage during saving by changing `maximum_memory_usage`. The default is `model.save_pretrained(..., maximum_memory_usage = 0.75)`. Reduce it to say 0.5 to use 50% of GPU peak memory or lower. This can reduce OOM crashes during saving. ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-do-i-manually-save-to-gguf) How do I manually save to GGUF? First save your model to 16bit via: Compile llama.cpp from source like below: Then, save the model to F16: ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#why-is-q8_k_xl-slower-than-q8_0-gguf) Why is Q8\_K\_XL slower than Q8\_0 GGUF? On Mac devices, it seems like that BF16 might be slower than F16. Q8\_K\_XL upcasts some layers to BF16, so hence the slowdown, We are actively changing our conversion process to make F16 the default choice for Q8\_K\_XL to reduce performance hits. ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-to-do-evaluation) How to do Evaluation To set up evaluation in your training run, you first have to split your dataset into a training and test split. You should **always shuffle the selection of the dataset**, otherwise your evaluation is wrong! Then, we can set the training arguments to enable evaluation. Reminder evaluation can be very very slow especially if you set `eval_steps = 1` which means you are evaluating every single step. If you are, try reducing the eval\_dataset size to say 100 rows or something. ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#evaluation-loop-out-of-memory-or-crashing) Evaluation Loop - Out of Memory or crashing. A common issue when you OOM is because you set your batch size too high. Set it lower than 2 to use less VRAM. Also use `fp16_full_eval=True` to use float16 for evaluation which cuts memory by 1/2. First split your training dataset into a train and test split. Set the trainer settings for evaluation to: This will cause no OOMs and make it somewhat faster. You can also use `bf16_full_eval=True` for bf16 machines. By default Unsloth should have set these flags on by default as of June 2025. ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-do-i-do-early-stopping) How do I do Early Stopping? If you want to stop the finetuning / training run since the evaluation loss is not decreasing, then you can use early stopping which stops the training process. Use `EarlyStoppingCallback`. As usual, set up your trainer and your evaluation dataset. The below is used to stop the training run if the `eval_loss` (the evaluation loss) is not decreasing after 3 steps or so. We then add the callback which can also be customized: Then train the model as usual via `trainer.train() .` ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#downloading-gets-stuck-at-90-to-95) Downloading gets stuck at 90 to 95% If your model gets stuck at 90, 95% for a long time before you can disable some fast downloading processes to force downloads to be synchronous and to print out more error messages. Simply use `UNSLOTH_STABLE_DOWNLOADS=1` before any Unsloth import. ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#runtimeerror-cuda-error-device-side-assert-triggered) RuntimeError: CUDA error: device-side assert triggered Restart and run all, but place this at the start before any Unsloth import. Also please file a bug report asap thank you! ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#all-labels-in-your-dataset-are-100.-training-losses-will-be-all-0) All labels in your dataset are -100. Training losses will be all 0. This means that your usage of `train_on_responses_only` is incorrect for that particular model. train\_on\_responses\_only allows you to mask the user question, and train your model to output the assistant response with higher weighting. This is known to increase accuracy by 1% or more. See our [**LoRA Hyperparameters Guide**](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide) for more details. For Llama 3.1, 3.2, 3.3 type models, please use the below: For Gemma 2, 3. 3n models, use the below: ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#unsloth-is-slower-than-expected) Unsloth is slower than expected? If your speed seems slower at first, it’s likely because `torch.compile` typically takes ~5 minutes (or longer) to warm up and finish compiling. Make sure you measure throughput **after** it’s fully loaded as over longer runs, Unsloth should be much faster. To disable use: ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#some-weights-of-gemma3nforconditionalgeneration-were-not-initialized-from-the-model-checkpoint) Some weights of Gemma3nForConditionalGeneration were not initialized from the model checkpoint This is a critical error, since this means some weights are not parsed correctly, which will cause incorrect outputs. This can normally be fixed by upgrading Unsloth `pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth_zoo` Then upgrade transformers and timm: `pip install --upgrade --force-reinstall --no-cache-dir --no-deps transformers timm` However if the issue still persists, please file a bug report asap! ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#notimplementederror-a-utf-8-locale-is-required.-got-ansi) NotImplementedError: A UTF-8 locale is required. Got ANSI See https://github.com/googlecolab/colabtools/issues/3409 In a new cell, run the below: ### [](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#citing-unsloth) Citing Unsloth If you are citing the usage of our model uploads, use the below Bibtex. This is for Qwen3-30B-A3B-GGUF Q8\_K\_XL: To cite the usage of our Github package or our work in general: [PreviousVision Fine-tuning](https://unsloth.ai/docs/basics/vision-fine-tuning) [NextHugging Face Hub, XET debugging](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging) Last updated 2 months ago Was this helpful? * [Fine-tuning a new model not supported by Unsloth?](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#fine-tuning-a-new-model-not-supported-by-unsloth) * [Running in Unsloth works well, but after exporting & running on other platforms, the results are poor](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#running-in-unsloth-works-well-but-after-exporting-and-running-on-other-platforms-the-results-are-poo) * [Saving to GGUF / vLLM 16bit crashes](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#saving-to-gguf-vllm-16bit-crashes) * [How do I manually save to GGUF?](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-do-i-manually-save-to-gguf) * [Why is Q8\_K\_XL slower than Q8\_0 GGUF?](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#why-is-q8_k_xl-slower-than-q8_0-gguf) * [How to do Evaluation](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-to-do-evaluation) * [Evaluation Loop - Out of Memory or crashing.](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#evaluation-loop-out-of-memory-or-crashing) * [How do I do Early Stopping?](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#how-do-i-do-early-stopping) * [Downloading gets stuck at 90 to 95%](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#downloading-gets-stuck-at-90-to-95) * [RuntimeError: CUDA error: device-side assert triggered](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#runtimeerror-cuda-error-device-side-assert-triggered) * [All labels in your dataset are -100. Training losses will be all 0.](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#all-labels-in-your-dataset-are-100.-training-losses-will-be-all-0) * [Unsloth is slower than expected?](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#unsloth-is-slower-than-expected) * [Some weights of Gemma3nForConditionalGeneration were not initialized from the model checkpoint](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#some-weights-of-gemma3nforconditionalgeneration-were-not-initialized-from-the-model-checkpoint) * [NotImplementedError: A UTF-8 locale is required. Got ANSI](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#notimplementederror-a-utf-8-locale-is-required.-got-ansi) * [Citing Unsloth](https://unsloth.ai/docs/basics/troubleshooting-and-faqs#citing-unsloth) Was this helpful? Copy from huggingface_hub import snapshot_download snapshot_download("unsloth/DeepSeek-OCR", local_dir = "deepseek_ocr") model, tokenizer = FastVisionModel.from_pretrained( "./deepseek_ocr", load_in_4bit = False, # Use 4bit to reduce memory use. False for 16bit LoRA. auto_model = AutoModel, trust_remote_code = True, # Enable to support new models unsloth_force_compile = True, use_gradient_checkpointing = "unsloth", # True or "unsloth" for long context ) Copy model.save_pretrained_merged("merged_model", tokenizer, save_method = "merged_16bit",) Copy apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \ -DBUILD_SHARED_LIBS=ON -DGGML_CUDA=ON -DLLAMA_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-* llama.cpp Copy python llama.cpp/convert_hf_to_gguf.py merged_model \ --outfile model-F16.gguf --outtype f16 \ --split-max-size 50G Copy # For BF16: python llama.cpp/convert_hf_to_gguf.py merged_model \ --outfile model-BF16.gguf --outtype bf16 \ --split-max-size 50G # For Q8_0: python llama.cpp/convert_hf_to_gguf.py merged_model \ --outfile model-Q8_0.gguf --outtype q8_0 \ --split-max-size 50G Copy new_dataset = dataset.train_test_split( test_size = 0.01, # 1% for test size can also be an integer for # of rows shuffle = True, # Should always set to True! seed = 3407, ) train_dataset = new_dataset["train"] # Dataset for training eval_dataset = new_dataset["test"] # Dataset for evaluation Copy from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( args = SFTConfig( fp16_full_eval = True, # Set this to reduce memory usage per_device_eval_batch_size = 2,# Increasing this will use more memory eval_accumulation_steps = 4, # You can increase this include of batch_size eval_strategy = "steps", # Runs eval every few steps or epochs. eval_steps = 1, # How many evaluations done per # of training steps ), train_dataset = new_dataset["train"], eval_dataset = new_dataset["test"], ... ) trainer.train() Copy new_dataset = dataset.train_test_split(test_size = 0.01) from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( args = SFTConfig( fp16_full_eval = True, per_device_eval_batch_size = 2, eval_accumulation_steps = 4, eval_strategy = "steps", eval_steps = 1, ), train_dataset = new_dataset["train"], eval_dataset = new_dataset["test"], ... ) Copy from trl import SFTConfig, SFTTrainer trainer = SFTTrainer( args = SFTConfig( fp16_full_eval = True, per_device_eval_batch_size = 2, eval_accumulation_steps = 4, output_dir = "training_checkpoints", # location of saved checkpoints for early stopping save_strategy = "steps", # save model every N steps save_steps = 10, # how many steps until we save the model save_total_limit = 3, # keep only 3 saved checkpoints to save disk space eval_strategy = "steps", # evaluate every N steps eval_steps = 10, # how many steps until we do evaluation load_best_model_at_end = True, # MUST USE for early stopping metric_for_best_model = "eval_loss", # metric we want to early stop on greater_is_better = False, # the lower the eval loss, the better ), model = model, tokenizer = tokenizer, train_dataset = new_dataset["train"], eval_dataset = new_dataset["test"], ) Copy from transformers import EarlyStoppingCallback early_stopping_callback = EarlyStoppingCallback( early_stopping_patience = 3, # How many steps we will wait if the eval loss doesn't decrease # For example the loss might increase, but decrease after 3 steps early_stopping_threshold = 0.0, # Can set higher - sets how much loss should decrease by until # we consider early stopping. For eg 0.01 means if loss was # 0.02 then 0.01, we consider to early stop the run. ) trainer.add_callback(early_stopping_callback) Copy import os os.environ["UNSLOTH_STABLE_DOWNLOADS"] = "1" from unsloth import FastLanguageModel Copy import os os.environ["UNSLOTH_COMPILE_DISABLE"] = "1" os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1" Copy from unsloth.chat_templates import train_on_responses_only trainer = train_on_responses_only( trainer, instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n", response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n", ) Copy from unsloth.chat_templates import train_on_responses_only trainer = train_on_responses_only( trainer, instruction_part = "user\n", response_part = "model\n", ) Copy import os os.environ["UNSLOTH_COMPILE_DISABLE"] = "1" Copy import locale locale.getpreferredencoding = lambda: "UTF-8" Copy @misc{unsloth_2025_qwen3_30b_a3b, author = {Unsloth AI and Han-Chen, Daniel and Han-Chen, Michael}, title = {Qwen3-30B-A3B-GGUF:Q8\_K\_XL}, year = {2025}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/unsloth/Qwen3-30B-A3B-GGUF}} } Copy @misc{unsloth, author = {Unsloth AI and Han-Chen, Daniel and Han-Chen, Michael}, title = {Unsloth}, year = {2025}, publisher = {Github}, howpublished = {\url{https://github.com/unslothai/unsloth}} } --- # Chat Templates | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/basics/chat-templates.md) . In our GitHub, we have a list of every chat template Unsloth uses including for Llama, Mistral, Phi-4 etc. So if you need any pointers on the formatting or use case, you can view them here: [github.com/unslothai/unsloth/blob/main/unsloth/chat\_templates.py](https://github.com/unslothai/unsloth/blob/main/unsloth/chat_templates.py) #### [](https://unsloth.ai/docs/basics/chat-templates#list-of-colab-chat-template-notebooks) List of Colab chat template notebooks: * [Conversational](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(1B_and_3B)-Conversational.ipynb) * [ChatML](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3_(8B)-Ollama.ipynb) * [Ollama](https://colab.research.google.com/drive/1WZDi7APtQ9VsvOrQSSC5DDtxq159j8iZ?usp=sharing) * [Text Classification](https://github.com/timothelaborie/text_classification_scripts/blob/main/unsloth_classification.ipynb) by Timotheeee * [Multiple Datasets](https://colab.research.google.com/drive/1njCCbE1YVal9xC83hjdo2hiGItpY_D6t?usp=sharing) by Flail ### [](https://unsloth.ai/docs/basics/chat-templates#adding-new-tokens) Adding new tokens Unsloth has a function called `add_new_tokens` which allows you to add new tokens to your finetune. For example if you want to add ``, `` and `` we can do the following: Copy model, tokenizer = FastLanguageModel.from_pretrained(...) from unsloth import add_new_tokens add_new_tokens(model, tokenizer, new_tokens = ["", "", ""]) model = FastLanguageModel.get_peft_model(...) Note - you MUST always call `add_new_tokens` before `FastLanguageModel.get_peft_model`! [](https://unsloth.ai/docs/basics/chat-templates#multi-turn-conversations) Multi turn conversations -------------------------------------------------------------------------------------------------------- An issue if you didn't notice is the Alpaca dataset is single turn, whilst remember using ChatGPT was interactive and you can talk to it in multiple turns. For example, the left is what we want, but the right which is the Alpaca dataset only provides singular conversations. We want the finetuned language model to somehow learn how to do multi turn conversations just like ChatGPT. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-2a65cd74ddd03a6bcbbc9827d9d034e4879a8e6a%252Fdiff.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=d776164b&sv=2) So we introduced the `conversation_extension` parameter, which essentially selects some random rows in your single turn dataset, and merges them into 1 conversation! For example, if you set it to 3, we randomly select 3 rows and merge them into 1! Setting them too long can make training slower, but could make your chatbot and final finetune much better! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-2b1b3494b260f1102942d86143a885225c6a06f2%252Fcombine.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=f26d5bca&sv=2) Then set `output_column_name` to the prediction / output column. For the Alpaca dataset, it would be the output column. We then use the `standardize_sharegpt` function to just make the dataset in a correct format for finetuning! Always call this! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-7bf83bf802191bda9e417bbe45afa181e7f24f38%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=936b7eae&sv=2) [](https://unsloth.ai/docs/basics/chat-templates#customizable-chat-templates) Customizable Chat Templates -------------------------------------------------------------------------------------------------------------- We can now specify the chat template for finetuning itself. The very famous Alpaca format is below: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-59737e6dcb09fed15487d5a57c69f07cb40bb8e7%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=81f68351&sv=2) But remember we said this was a bad idea because ChatGPT style finetunes require only 1 prompt? Since we successfully merged all dataset columns into 1 using Unsloth, we essentially can create the below style chat template with 1 input column (instruction) and 1 output: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-d54582ae98c396d51bfb85628b46c54f2517d030%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=68b8c20&sv=2) We just require you must put a `{INPUT}` field for the instruction and an `{OUTPUT}` field for the model's output field. We in fact allow an optional `{SYSTEM}` field as well which is useful to customize a system prompt just like in ChatGPT. For example, below are some cool examples which you can customize the chat template to be: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-cc455dc380d3d44ef136e485754964159dc773d8%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=82deed73&sv=2) For the ChatML format used in OpenAI models: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-15bfca9cfadf10d54b4d3f66e3050044317d62c5%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=c0e648a7&sv=2) Or you can use the Llama-3 template itself (which only functions by using the instruct version of Llama-3): We in fact allow an optional `{SYSTEM}` field as well which is useful to customize a system prompt just like in ChatGPT. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-80a2ed4de2ca323ac192c513cac65e9e8bf475db%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=467af04b&sv=2) Or in the Titanic prediction task where you had to predict if a passenger died or survived in this Colab notebook which includes CSV and Excel uploading: [https://colab.research.google.com/drive/1VYkncZMfGFkeCEgN2IzbZIKEDkyQuJAS?usp=sharing](https://colab.research.google.com/drive/1VYkncZMfGFkeCEgN2IzbZIKEDkyQuJAS?usp=sharing) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-20911ab305c1a10e85859c703157b80175141eb1%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=957f7f78&sv=2) [](https://unsloth.ai/docs/basics/chat-templates#applying-chat-templates-with-unsloth) Applying Chat Templates with Unsloth -------------------------------------------------------------------------------------------------------------------------------- For datasets that usually follow the common chatml format, the process of preparing the dataset for training or finetuning, consists of four simple steps: * Check the chat templates that Unsloth currently supports:\\ This will print out the list of templates currently supported by Unsloth. Here is an example output:\\ \\ * Use `get_chat_template` to apply the right chat template to your tokenizer:\\ \\ * Define your formatting function. Here's an example:\\ This function loops through your dataset applying the chat template you defined to each sample.\\ * Finally, let's load the dataset and apply the required modifications to our dataset: \\ If your dataset uses the ShareGPT format with "from"/"value" keys instead of the ChatML "role"/"content" format, you can use the `standardize_sharegpt` function to convert it first. The revised code will now look as follows: \\ [](https://unsloth.ai/docs/basics/chat-templates#more-information) More Information ---------------------------------------------------------------------------------------- Assuming your dataset is a list of list of dictionaries like the below: You can use our `get_chat_template` to format it. Select `chat_template` to be any of `zephyr, chatml, mistral, llama, alpaca, vicuna, vicuna_old, unsloth`, and use `mapping` to map the dictionary values `from`, `value` etc. `map_eos_token` allows you to map `<|im_end|>` to EOS without any training. You can also make your own custom chat templates! For example our internal chat template we use is below. You must pass in a `tuple` of `(custom_template, eos_token)` where the `eos_token` must be used inside the template. [PreviousHugging Face Hub, XET debugging](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging) [NextContinued Pretraining](https://unsloth.ai/docs/basics/continued-pretraining) Last updated 2 months ago Was this helpful? * [Adding new tokens](https://unsloth.ai/docs/basics/chat-templates#adding-new-tokens) * [Multi turn conversations](https://unsloth.ai/docs/basics/chat-templates#multi-turn-conversations) * [Customizable Chat Templates](https://unsloth.ai/docs/basics/chat-templates#customizable-chat-templates) * [Applying Chat Templates with Unsloth](https://unsloth.ai/docs/basics/chat-templates#applying-chat-templates-with-unsloth) * [More Information](https://unsloth.ai/docs/basics/chat-templates#more-information) Was this helpful? Copy from unsloth.chat_templates import CHAT_TEMPLATES print(list(CHAT_TEMPLATES.keys())) Copy ['unsloth', 'zephyr', 'chatml', 'mistral', 'llama', 'vicuna', 'vicuna_old', 'vicuna old', 'alpaca', 'gemma', 'gemma_chatml', 'gemma2', 'gemma2_chatml', 'llama-3', 'llama3', 'phi-3', 'phi-35', 'phi-3.5', 'llama-3.1', 'llama-31', 'llama-3.2', 'llama-3.3', 'llama-32', 'llama-33', 'qwen-2.5', 'qwen-25', 'qwen25', 'qwen2.5', 'phi-4', 'gemma-3', 'gemma3'] Copy from unsloth.chat_templates import get_chat_template tokenizer = get_chat_template( tokenizer, chat_template = "gemma-3", # change this to the right chat_template name ) Copy def formatting_prompts_func(examples): convos = examples["conversations"] texts = [tokenizer.apply_chat_template(convo, tokenize = False, add_generation_prompt = False) for convo in convos] return { "text" : texts, } Copy # Import and load dataset from datasets import load_dataset dataset = load_dataset("repo_name/dataset_name", split = "train") # Apply the formatting function to your dataset using the map method dataset = dataset.map(formatting_prompts_func, batched = True,) Copy # Import dataset from datasets import load_dataset dataset = load_dataset("mlabonne/FineTome-100k", split = "train") # Convert your dataset to the "role"/"content" format if necessary from unsloth.chat_templates import standardize_sharegpt dataset = standardize_sharegpt(dataset) # Apply the formatting function to your dataset using the map method dataset = dataset.map(formatting_prompts_func, batched = True,) Copy [\ [{'from': 'human', 'value': 'Hi there!'},\ {'from': 'gpt', 'value': 'Hi how can I help?'},\ {'from': 'human', 'value': 'What is 2+2?'}],\ [{'from': 'human', 'value': 'What's your name?'},\ {'from': 'gpt', 'value': 'I'm Daniel!'},\ {'from': 'human', 'value': 'Ok! Nice!'},\ {'from': 'gpt', 'value': 'What can I do for you?'},\ {'from': 'human', 'value': 'Oh nothing :)'},],\ ] Copy from unsloth.chat_templates import get_chat_template tokenizer = get_chat_template( tokenizer, chat_template = "chatml", # Supports zephyr, chatml, mistral, llama, alpaca, vicuna, vicuna_old, unsloth mapping = {"role" : "from", "content" : "value", "user" : "human", "assistant" : "gpt"}, # ShareGPT style map_eos_token = True, # Maps <|im_end|> to instead ) def formatting_prompts_func(examples): convos = examples["conversations"] texts = [tokenizer.apply_chat_template(convo, tokenize = False, add_generation_prompt = False) for convo in convos] return { "text" : texts, } pass from datasets import load_dataset dataset = load_dataset("philschmid/guanaco-sharegpt-style", split = "train") dataset = dataset.map(formatting_prompts_func, batched = True,) Copy unsloth_template = \ "{{ bos_token }}"\ "{{ 'You are a helpful assistant to the user\n' }}"\ ""\ "
"\ "
"\ "{{ '>>> User: ' + message['content'] + '\n' }}"\ "
"\ "{{ '>>> Assistant: ' + message['content'] + eos_token + '\n' }}"\ "
"\ "
"\ "
"\ "{{ '>>> Assistant: ' }}"\ "
" unsloth_eos_token = "eos_token" tokenizer = get_chat_template( tokenizer, chat_template = (unsloth_template, unsloth_eos_token,), # You must provide a template and EOS token mapping = {"role" : "from", "content" : "value", "user" : "human", "assistant" : "gpt"}, # ShareGPT style map_eos_token = True, # Maps <|im_end|> to instead ) --- # How to Run Local AI Models with Hermes Agent | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/hermes-agent.md) . This guide enables you to run open LLMs locally with **Hermes Agent** via [**Unsloth**](https://github.com/unslothai/unsloth) . Hermes Agent by Nous Research is an **open-source** autonomous AI agent that connects to a model endpoint, executes tasks, and improves over time through memory and learned skills. Hermes will work with any **local model** exposed through Unsloth’s **OpenAI-compatible API**, including: DeepSeek, Qwen, Gemma, and more. Hermes acts as the agent client, while Unsloth loads and serves models via the [local API](https://unsloth.ai/docs/basics/api) entirely offline. After setup, every prompt sent through Hermes will run using your local model on your device. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FYhiTzkyz0fVPMxIfWg6D%252FScreenshot_20260718_122219.png%3Falt%3Dmedia%26token%3Da8063d58-0d49-467a-aa2a-35b30c448390&width=768&dpr=3&quality=100&sign=39ad9365&sv=2) Qwen3.5 running locally in Hermes via Unsloth. [Setup Hermes](https://unsloth.ai/docs/integrations/hermes-agent#setup-hermes-agent) [🦥 Connect your local model](https://unsloth.ai/docs/integrations/hermes-agent#integrate-hermes-with-unsloth-api) In this tutorial, you’ll install Hermes and configure it to use `unsloth/Qwen3.6-27B-GGUF` served from Unsloth. Prefer a different model? Swap in any other model by loading it in Unsloth and updating the configuration. ### [](https://unsloth.ai/docs/integrations/hermes-agent#setup-hermes-agent) Setup Hermes Agent **Prerequisites:** The [Hermes](https://github.com/NousResearch/hermes-agent/blob/main/website/docs/getting-started/installation.md) command-line installer supports Linux, macOS, and WSL2. Make sure **Git** is installed; on Linux, also install **curl** and **xz-utils**. The installer automatically provisions `uv`, Python 3.11, Node.js 22, `ripgrep`, and `ffmpeg`. #### [](https://unsloth.ai/docs/integrations/hermes-agent#id-1.-run-the-installer) 1\. Run the installer Copy curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash The installer: * Detects your platform and checks dependencies. * Clones Hermes to `~/.hermes/hermes-agent/`. * Creates a Python virtual environment and installs the Python dependencies. * Installs the browser-tool dependencies and Playwright’s Chromium engine. * Adds the `hermes` command and launches the setup wizard. Playwright may request `sudo` to install Chromium’s shared system libraries. Hermes itself does not require root access. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FnabBNaHLvb1rH58cQwyC%252FScreenshot%25202026-07-18%2520at%25205.31.57%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D4ac35864-89aa-45ce-84db-ad5b17f8b3e5&width=768&dpr=3&quality=100&sign=4348f46b&sv=2) #### [](https://unsloth.ai/docs/integrations/hermes-agent#id-2.-reload-your-shell-so-the-hermes-command-is-on-your-path) **2\. Reload your shell** so the `hermes` command is on your `PATH`: #### [](https://unsloth.ai/docs/integrations/hermes-agent#id-3.-verify-the-install) **3\. Verify the install:** If the command resolves, Hermes is installed. Everything lives under `~/.hermes/`: Path What it is `~/.hermes/config.yaml` Main settings (model, provider, tools, TTS, …) `~/.hermes/.env` API keys and other secrets `~/.hermes/hermes-agent/` The Hermes source + virtualenv `~/.hermes/cron/`, `sessions/`, `logs/` Runtime data `~/.hermes/skills/` Installed skills (synced from the Skills Hub) Full install reference: [hermes-agent.nousresearch.com/docs/getting-started/installation](https://hermes-agent.nousresearch.com/docs/getting-started/installation) . If the installer reports a missing prerequisite, install it and re-run the one-liner. The installer is idempotent. ### [](https://unsloth.ai/docs/integrations/hermes-agent#quickstart) ⚡ Quickstart After installing Hermes, we'll need to install Unsloth Studio to enable Hermes to serve and run inference of local models. 1. **Install or update Unsloth Studio.** Earlier versions don't expose the external API. See Installation. 2. **Launch Unsloth.** Note the port it starts on is usually `8000` or `8888`. You'll see it in the terminal output and in the browser URL (`http://localhost:PORT`). 3. **Load a model.** Click **New Chat**, pick or search a model (GGUF), and wait for it to finish loading. 4. **Connect Hermes.** Run `unsloth start hermes`. It mints an API key, writes the config, and launches Hermes against your loaded model. ### [](https://unsloth.ai/docs/integrations/hermes-agent#run-hermes-agent-with-unsloth-start) ⚡ Run Hermes Agent with `unsloth start` To launch Hermes directly with a model, run: With a model loaded in Unsloth Studio, run: ![Hermes Agent connected to a local model through Unsloth Studio](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FOaFdKwRASHSoxDQx06l5%252FScreenshot_20260714_152023.png%3Falt%3Dmedia%26token%3Dc7e13fb2-e03e-49d1-8efa-f7494578943a&width=768&dpr=3&quality=100&sign=d06367c5&sv=2) Hermes Agent running through its Unsloth Studio provider. Unsloth launches Hermes from a separate managed home with the Unsloth provider, model, and context settings already configured. Your existing Hermes setup is left unchanged. This managed home is temporary by default. To keep your sessions and state, add `--persist` from your first launch: To return to your latest session later, run: To reopen a specific session, use `--resume ` instead. See the complete [unsloth start](https://unsloth.ai/docs/integrations/unsloth-start) reference for model selection, remote connections, and advanced options. The setup wizard below remains available if you prefer to manage the Hermes provider yourself. ### [](https://unsloth.ai/docs/integrations/hermes-agent#creating-an-api-key) 🔑 Creating an API key 1. Open the sidebar, click your **Unsloth** avatar at the bottom-left. 2. Go to **Settings** → **API**. 3. Enter a friendly name (e.g. `hermes-agent-macbook`). 4. _(Optional)_ Set an expiry. 5. Click **Create**. 6. **Copy the key immediately.** Unsloth stores only a hash and you won't be able to view it again. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FrWnqY5PgvZdbWK8pmlD5%252Fimage.png%3Falt%3Dmedia%26token%3D3507e7d7-b552-447f-b503-2418802b8f6e&width=768&dpr=3&quality=100&sign=d4395e2c&sv=2) All keys start with the `sk-unsloth-` prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with `401 Unauthorized`. ### [](https://unsloth.ai/docs/integrations/hermes-agent#integrate-hermes-with-unsloth-api) 🦥 Integrate Hermes with Unsloth API Hermes sends each chat turn to a configured inference provider and connects to **OpenAI-compatible** endpoints. Configure the provider during installation or later in the setup wizard. **1\. Open the setup wizard:** Pick **Model & Provider** from the "What would you like to do?" menu to configure only the inference endpoint, or **Full Setup** to walk through everything (TTS, tools, messaging gateway, agent settings). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FChF6By1Dl44hVQl63kGW%252FScreenshot%25202026-07-18%2520at%25205.34.39%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D371a287e-1264-47d2-af3e-fa3ed302e038&width=768&dpr=3&quality=100&sign=46dd93af&sv=2) **2\. Select the Custom OpenAI-compatible endpoint** when Hermes prompts you for an inference provider. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FA5VVaVsL7GzZGq9Dn3Bt%252Fhermes_agent_setup_2.png%3Falt%3Dmedia%26token%3Da234d29d-14f6-431f-be1f-9a4e14f5951d&width=768&dpr=3&quality=100&sign=4a1b9d6&sv=2) **3\. Fill in the prompts** as Hermes walks through them: Prompt Value **API base URL** `http://localhost:8888/v1` _(your Unsloth port +_ `_/v1_`_)_ **API key** Your `sk-unsloth-…` key **Detected model: … Use this model?** `Y` _(Hermes auto-detects the model via_ `_GET /v1/models_`_)_ **Context length in tokens** _(leave blank for auto-detect)_ **Display name** Anything you like, e.g. `unsloth-api` Hermes verifies the endpoint against `/v1/models` and confirms the detected model before continuing. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FehQjrLJ0ztoSdNp94dfH%252Fhermes_agent_setup_3.png%3Falt%3Dmedia%26token%3D641c6e6a-44c3-4013-8c08-32472b256283&width=768&dpr=3&quality=100&sign=18ef18a3&sv=2) **4\. Accept defaults for the remaining prompts** (TTS, tools, messaging gateway, agent settings) you can reconfigure any of them later. Hermes writes everything to `~/.hermes/config.yaml` and `~/.hermes/.env`. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FWbj2YWKpXw9T2jR0QQjh%252Fhermes_agent_setup_5.png%3Falt%3Dmedia%26token%3D59985534-69e3-46ec-8ab9-1db8829c1721&width=768&dpr=3&quality=100&sign=5ec4f41a&sv=2) **5\. Launch Hermes:** The startup banner shows your Unsloth model name in the status bar (e.g. `unsloth/Qwen3.6-27B-GGUF`), and the prompt is ready for input. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FU8UlRyw9n1SrENvZXSjD%252Fhermes_unsloth_run.png%3Falt%3Dmedia%26token%3D6fd14920-a287-4aed-815e-33b7e4d09c68&width=768&dpr=3&quality=100&sign=d9b6132d&sv=2) To reconfigure just the model later, run `hermes setup model`. To edit the config file directly, `hermes config edit` opens `~/.hermes/config.yaml` in your `$EDITOR`. ### [](https://unsloth.ai/docs/integrations/hermes-agent#optional-tune-the-unsloth-server) Optional: tune the Unsloth server `unsloth run` starts the local API server and loads a model for your app to connect to. You can also customize how the server behaves when starting it. Use `--disable-tools` when driving Hermes (or any external agent with its own tools). By default Unsloth Studio runs its own server-side tools, which swallows the agent's tool calls, so Hermes answers but never runs its tools. `--disable-tools` switches to passthrough, so Hermes's own tools are used. Use `--reasoning off` to turn thinking off, or `--reasoning on` to turn it on for models that support reasoning. This starts the server on `0.0.0.0:8888`, allowing other devices on your local network to connect. `-p` changes which port the server runs on. If you want phones, laptops, or other devices on your network to connect to the API server, start it with `-H 0.0.0.0`. Some apps may still override generation settings for individual requests. For more advanced runtime configuration, see the main [API tuning](https://unsloth.ai/docs/basics/api#unsloth-run-command) section. [PreviousOpenRouter](https://unsloth.ai/docs/integrations/connections/openrouter) [NextOpenClaw](https://unsloth.ai/docs/integrations/openclaw) Last updated 3 days ago Was this helpful? * [Setup Hermes Agent](https://unsloth.ai/docs/integrations/hermes-agent#setup-hermes-agent) * [⚡ Quickstart](https://unsloth.ai/docs/integrations/hermes-agent#quickstart) * [⚡ Run Hermes Agent with unsloth start](https://unsloth.ai/docs/integrations/hermes-agent#run-hermes-agent-with-unsloth-start) * [🔑 Creating an API key](https://unsloth.ai/docs/integrations/hermes-agent#creating-an-api-key) * [🦥 Integrate Hermes with Unsloth API](https://unsloth.ai/docs/integrations/hermes-agent#integrate-hermes-with-unsloth-api) * [Optional: tune the Unsloth server](https://unsloth.ai/docs/integrations/hermes-agent#optional-tune-the-unsloth-server) Was this helpful? bash Copy source ~/.bashrc zsh Copy source ~/.zshrc Copy hermes --version Copy unsloth start hermes \ --model unsloth/gemma-4-E2B-it-GGUF:UD-Q4_K_XL \ --context-length 32768 Copy unsloth start hermes Copy unsloth start hermes --persist Copy unsloth start hermes --persist --continue Copy hermes setup Copy hermes Copy # Serve Hermes (--disable-tools passes the agent's own tools through) unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ --disable-tools \ --reasoning off \ -p 8888 Copy # Expose the API on your local network unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ -H 0.0.0.0 \ -p 8888 --- # Connect OpenAI to Unsloth: Run GPT Models in Local Chat | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/openai.md) . Learn how to connect OpenAI models, including GPT-5.5 to [Unsloth](https://github.com/unslothai/unsloth) so you can chat with all of them in an open-source local UI chat interface. By connecting your OpenAI API key, you can run GPT models inside Unsloth with features like [web search](https://unsloth.ai/docs/integrations/connections/openai#web-search-and-thinking) , tool-calling, [code execution](https://unsloth.ai/docs/integrations/connections/openai#code-execution) , [image generation](https://unsloth.ai/docs/integrations/connections/openai#image-generation) , reusable code containers, and [prompt caching](https://unsloth.ai/docs/integrations/connections/openai#prompt-caching) . This guide walks you through creating an OpenAI API key, connecting OpenAI as a provider, loading available models, and troubleshooting common setup issues. ### [](https://unsloth.ai/docs/integrations/connections/openai#setup) Setup 1 #### [](https://unsloth.ai/docs/integrations/connections/openai#create-an-openai-api-key) Create an OpenAI API key Create an API key from the [OpenAI dashboard](https://platform.openai.com/api-keys) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FNbLWLZa1OBYQrwX0mQVv%252Funsloth_openai_api.gif%3Falt%3Dmedia%26token%3D28187d50-8638-4e60-ba33-729efd13f1c1&width=768&dpr=3&quality=100&sign=6ac249d2&sv=2) 2 #### [](https://unsloth.ai/docs/integrations/connections/openai#configure-connections) Configure Connections Next, connect your provider to Unsloth. 1. Open **Settings** → **Connections**, then click **Add Connection.** 2. Select the OpenAI, then paste the API key you copied earlier. 3. Click **Reload Models** to refresh the list with models available to your account. 4. Choose the models you want to enable, then hit save. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FlAXYV1wS2lxeS36Lo8Dj%252Fexport-1779457818713.gif%3Falt%3Dmedia%26token%3D99c6f0b2-2fb1-40e4-a15a-325bcef8c2f7&width=768&dpr=3&quality=100&sign=c6be9fd3&sv=2) 3 #### [](https://unsloth.ai/docs/integrations/connections/openai#ready-to-chat) Ready to Chat The models you enabled will now appear under Connected in the Select Model dropdown. Supported GPT models can expose extra controls including image generation, thinking, web search and code execution. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FgKCJ3zsZzZuDqBAigsfG%252FScreenshot%25202026-05-26%2520at%25201.31.03%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dc8df4285-ba69-4091-8978-e4197f6a12d2&width=768&dpr=3&quality=100&sign=d8a5d0b3&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/openai#code-execution) Code Execution When enabled, supported OpenAI models can run code in a provider sandbox to solve problems, analyse data, and work with files. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fn65DzRsTGrflcqJBv67y%252FScreenshot%25202026-05-26%2520at%25206.11.54%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Df620e467-2202-4843-ae16-e169a89f0a57&width=768&dpr=3&quality=100&sign=21f105f3&sv=2) OpenAI uses reusable shell containers. In **Code Execution** settings, you can set the idle timeout, create containers, select the active container, refresh the list, or delete old containers. Select the same container in a new thread to continue with its files and state. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FBR0zCMzeC4IwPSmlx12A%252Fimage.png%3Falt%3Dmedia%26token%3D3d9e80ee-4ceb-4454-9620-165650a29c0a&width=768&dpr=3&quality=100&sign=3783afda&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/openai#prompt-caching) Prompt Caching Prompt caching reduces latency and cost when requests reuse the same long prefix. It is supported for compatible providers and servers, including OpenAI. Use the **Prompt caching** setting in the side panel to control caching behaviour for supported connections. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FKA6iU3qCFhUiq0KI16aU%252FScreenshot%25202026-05-26%2520at%25203.28.44%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D41a43794-792d-48f4-8774-8b9d85702dfc&width=768&dpr=3&quality=100&sign=3378c69e&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fq4csmTX5rS9isMkNkhTX%252FPrompt%2520Caching%2520Diagram%2520%281%29.png%3Falt%3Dmedia%26token%3Dade433bf-5eaf-4146-a266-525a85a6c98d&width=768&dpr=3&quality=100&sign=42ba571a&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/openai#web-search-and-thinking) Web Search & Thinking Provider-side web search is available for supported models from OpenAI. The **Think** control adapts to the selected model: some models use an on/off toggle, while reasoning-effort models use model specific thinking levels. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FC0Ed4kzN9h6c0ogEn6NT%252Fwebsearch%2520api.png%3Falt%3Dmedia%26token%3Dd3335222-1d9e-4021-9bf9-734c6acf1fc0&width=768&dpr=3&quality=100&sign=ecb70969&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/openai#image-generation) Image Generation Just like GPT, Unsloth also supports image generation. You can directly edit an image by clicking the “Edit Image” button and entering a new prompt to refine or regenerate it. Images are generated automatically when requested, but you can toggle this behavior off. A download button is also available, allowing you to save the image in its original full resolution. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F5ZLU2BY19MN4yiYLx2Br%252FScreenshot%25202026-05-26%2520at%25206.04.34%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D44327a8f-c14d-4541-9a02-f19f375b4326&width=768&dpr=3&quality=100&sign=44d8a39e&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FDQ1bJ5KF8CiyItrfOrIt%252FScreenshot%25202026-05-26%2520at%25206.06.07%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dff6b00bc-0a2f-4375-a454-49429cf839f2&width=768&dpr=3&quality=100&sign=a122ab34&sv=2) ### [](https://unsloth.ai/docs/integrations/connections/openai#troubleshooting) Troubleshooting If OpenAI fails to connect, check that the API key is valid and belongs to the correct OpenAI account. If a model does not appear after clicking **Load Models**, it may not be available for your account. You can enter the model ID manually or choose another model. [PreviousConnect a Provider](https://unsloth.ai/docs/integrations/connections) [NextAnthropic (Claude)](https://unsloth.ai/docs/integrations/connections/anthropic-claude) Last updated 1 month ago Was this helpful? * [Setup](https://unsloth.ai/docs/integrations/connections/openai#setup) * [Code Execution](https://unsloth.ai/docs/integrations/connections/openai#code-execution) * [Prompt Caching](https://unsloth.ai/docs/integrations/connections/openai#prompt-caching) * [Web Search & Thinking](https://unsloth.ai/docs/integrations/connections/openai#web-search-and-thinking) * [Image Generation](https://unsloth.ai/docs/integrations/connections/openai#image-generation) * [Troubleshooting](https://unsloth.ai/docs/integrations/connections/openai#troubleshooting) Was this helpful? --- # Fine-tuning LLMs with Blackwell, RTX 50 series & Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md) . Unsloth now supports NVIDIA’s Blackwell architecture GPUs, including RTX 50-series GPUs (5060–5090), RTX PRO 6000, and GPUS such as B200, B40, GB100, GB102 and more! You can read the official [NVIDIA blogpost here](https://developer.nvidia.com/blog/train-an-llm-on-an-nvidia-blackwell-desktop-with-unsloth-and-scale-it/) . Unsloth is now compatible with every NVIDIA GPU from 2018+ including the [DGX Spark](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth) . > **Our new** [**Docker image**](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#docker) > **supports Blackwell. Run the Docker image and start training!** [**Guide**](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth) ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#pip-install) Pip install Simply install Unsloth: Copy pip install unsloth If you see issues, another option is to create a separate isolated environment: Copy python -m venv unsloth source unsloth/bin/activate pip install unsloth Note it might be `pip3` or `pip3.13` and also `python3` or `python3.13` You might encounter some Xformers issues, in which cause you should build from source: Copy # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH_CUDA_ARCH_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#docker) Docker [`**unsloth/unsloth**`](https://hub.docker.com/r/unsloth/unsloth) is Unsloth's only Docker image. For Blackwell and 50-series GPUs, use this same image - no separate image needed. For installation instructions, please follow our [Unsloth Docker guide](https://unsloth.ai/docs/blog/how-to-fine-tune-llms-with-unsloth-and-docker) . ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#uv) uv #### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#uv-advanced) uv (Advanced) The installation order is important, since we want the overwrite bundled dependencies with specific versions (namely, `xformers` and `triton`). 1. I prefer to use `uv` over `pip` as it's faster and better for resolving dependencies, especially for libraries which depend on `torch` but for which a specific `CUDA` version is required per this scenario. Install `uv` Create a project dir and venv: 2. Install `vllm` Note that we have to specify `cu128`, otherwise `vllm` will install `torch==2.7.0` but with `cu126`. 3. Install `unsloth` dependencies If you notice weird resolving issues due to Xformers, you can also install Unsloth from source without Xformers: 4. Download and build `xformers` (Optional) Xformers is optional, but it is definitely faster and uses less memory. We'll use PyTorch's native SDPA if you do not want Xformers. Building Xformers from source might be slow, so beware! Note that we have to explicitly set `TORCH_CUDA_ARCH_LIST=12.0`. 5. `transformers` Install any transformers version, but best to get the latest. ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#conda-or-mamba-advanced) Conda or mamba (Advanced) 1. Install `conda/mamba` Run the installation script Create a conda or mamba environment Activate newly created environment 2. Install `vllm` Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your `(unsloth-blackwell)user@machine:` Note that we have to specify `cu128`, otherwise `vllm` will install `torch==2.7.0` but with `cu126`. 3. Install `unsloth` dependencies Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your `(unsloth-blackwell)user@machine:` 4. Download and build `xformers` (Optional) Xformers is optional, but it is definitely faster and uses less memory. We'll use PyTorch's native SDPA if you do not want Xformers. Building Xformers from source might be slow, so beware! You should see the name of your environment as a prefix to your terminal shell like this your `(unsloth-blackwell)user@machine:` Note that we have to explicitly set `TORCH_CUDA_ARCH_LIST=12.0`. 5. Update `triton` Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your `(unsloth-blackwell)user@machine:` `triton>=3.3.1` is required for `Blackwell` support. 6. `Transformers` Install any transformers version, but best to get the latest. If you are using mamba as your package just replace conda with mamba for all commands shown above. ### [](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#wsl-specific-notes) WSL-Specific Notes If you're using WSL (Windows Subsystem for Linux) and encounter issues during xformers compilation (reminder Xformers is optional, but faster for training) follow these additional steps: 1. **Increase WSL Memory Limit** Create or edit the WSL configuration file: After making these changes, restart WSL: 2. **Install xformers** Use the following command to install xformers with optimized compilation for WSL: The `--no-build-isolation` flag helps avoid potential build issues in WSL environments. [PreviousDGX Spark and Unsloth](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth) Last updated 2 months ago Was this helpful? * [Pip install](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#pip-install) * [Docker](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#docker) * [uv](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#uv) * [Conda or mamba (Advanced)](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#conda-or-mamba-advanced) * [WSL-Specific Notes](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth#wsl-specific-notes) Was this helpful? Copy uv pip install unsloth Copy curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env Copy mkdir 'unsloth-blackwell' && cd 'unsloth-blackwell' uv venv .venv --python=3.12 --seed source .venv/bin/activate Copy uv pip install -U vllm --torch-backend=cu128 Copy uv pip install unsloth unsloth_zoo bitsandbytes Copy uv pip install -qqq \ "unsloth_zoo[base] @ git+https://github.com/unslothai/unsloth-zoo" \ "unsloth[base] @ git+https://github.com/unslothai/unsloth" Copy # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH_CUDA_ARCH_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. Copy uv pip install -U transformers Copy curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh" Copy bash Miniforge3-$(uname)-$(uname -m).sh Copy conda create --name unsloth-blackwell python==3.12 -y Copy conda activate unsloth-blackwell Copy pip install -U vllm --extra-index-url https://download.pytorch.org/whl/cu128 Copy pip install unsloth unsloth_zoo bitsandbytes Copy # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH_CUDA_ARCH_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. Copy pip install -U "triton>=3.3.1" Copy uv pip install -U transformers Copy # Create or edit .wslconfig in your Windows user directory # (typically C:\Users\YourUsername\.wslconfig) # Add these lines to the file [wsl2] memory=16GB # Minimum 16GB recommended for xformers compilation processors=4 # Adjust based on your CPU cores swap=2GB localhostForwarding=true Copy wsl --shutdown Copy # Set CUDA architecture for Blackwell GPUs export TORCH_CUDA_ARCH_LIST="12.0" # Install xformers from source with optimized build flags pip install -v --no-build-isolation -U git+https://github.com/facebookresearch/xformers.git@main#egg=xformers --- # 3x Faster LLM Training with Unsloth Kernels + Packing | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/blog/3x-faster-training-packing.md) . Unsloth now supports up to **5× faster** (typically 3x) training with our new custom **RoPE and MLP Triton kernels**, plus our new smart auto packing. Unsloth's new kernels + features not only increase training speed, but also further **reduces VRAM use (30% - 90%)** with no accuracy loss. [Unsloth GitHub](https://github.com/unslothai/unsloth) This means you can now train LLMs like [Qwen3](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune) \-4B not only on just **3GB VRAM**, but also 3x faster. Our auto [**padding-free**](https://unsloth.ai/docs/blog/3x-faster-training-packing#padding-free-by-default) uncontaminated packing is smartly enabled for all training runs without any changes, and all fast attention backends (FlashAttention 3, xFormers, SDPA). [Benchmarks](https://unsloth.ai/docs/blog/3x-faster-training-packing#analysis-and-benchmarks) show training losses match non-packing runs **exactly**. * **2.3x faster QK Rotary Embedding** fused Triton kernel with packing support * Updated SwiGLU, GeGLU kernels with **int64 indexing for long context** * **2.5x to 5x faster uncontaminated packing** with xformers, SDPA, FA3 backends * **2.1x faster padding free, 50% less VRAM**, 0% accuracy change * Unsloth also now has improved SFT loss stability and more predictable GPU utilization. * This new upgrade works **for all training methods** e.g. full fine-tuning, pretraining etc. ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#fused-qk-rope-triton-kernel-with-packing) 🥁Fused QK RoPE Triton Kernel with packing Back in December 2023, we introduced a RoPE kernel coded up in Triton as part of our Unsloth launch. In March 2024, a community member made end to end training 1-2% faster by optimizing the RoPE kernel to allow launching a block for a group of heads. See [PR 238](https://github.com/unslothai/unsloth/pull/238) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FewadBu05vK7zAmJRJcj6%252Frope_varlen_qk_rope_kernel_benchmark_v5.png%3Falt%3Dmedia%26token%3D04d277d4-c289-4943-9312-e3d3e2d60bec&width=768&dpr=3&quality=100&sign=3a1580af&sv=2) One issue is for each Q and K, there are 2 Triton kernels. We merged them into 1 Triton kernel now, and enabled variable length RoPE, which was imperative for padding free and packing support. This makes the RoPE kernel in micro benchmarks **2.3x faster on longer context lengths**, and 1.9x faster on shorter context lengths. We also eliminated all clones and contiguous transpose operations, and so **RoPE is now fully inplace**, reducing further GPU memory. Note for the backward pass, we see that `sin1 = -sin1` since: ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#int64-indexing-for-triton-kernels) 🚃Int64 Indexing for Triton Kernels During 500K long context training which we introduced in [500K Context Training](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning) , we would get CUDA out of bounds errors. This was because MLP kernels for SwiGLU, GeGLU had int32 indexing which is by default in Triton and CUDA. We can't just do `tl.program_id(0).to(tl.int64)` since training will be slightly slower due to int64 indexing. We instead make this a `LONG_INDEXING: tl.constexpr` variable so the Triton compiler can specialize this. This allows shorter and longer context runs to both run great! ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#why-is-padding-needed-and-mathematical-speedup) 🧮Why is padding needed & mathematical speedup Computers and GPUs cannot process different length datasets, so we have to pad them with 0s. This causes wastage. Assume we have a dataset of 50% short sequences S, and 50% long sequences L, then in the worst case, padding will cause token usage to be batchsize×L\\text{batchsize} \\times Lbatchsize×L since the longest sequence length dominates. By packing multiple examples into a single, long one-dimensional tensor, we can eliminate a significant amount of padding. In fact we get the below token usage: Token Usage\=batchsize2L+batchsize2S\\text{Token Usage} = \\frac{\\text{batchsize}}{2}L+\\frac{\\text{batchsize}}{2}SToken Usage\=2batchsize​L+2batchsize​S By some math and algebra, we can work out the speedup via: Speedup\=batchsize×Lbatchsize2L+batchsize2S\=2LL+S\\text{Speedup} = \\frac{\\text{batchsize} \\times L}{\\frac{\\text{batchsize}}{2}L+\\frac{\\text{batchsize}}{2}S} = 2 \\frac{L}{L + S}Speedup\=2batchsize​L+2batchsize​Sbatchsize×L​\=2L+SL​ By assuming S→0S\\rightarrow0S→0 then we get a 2x theoretical speedup since 2LL+0\=22 \\frac{L}{L + 0} = 22L+0L​\=2 By changing the ratio of 50% short sequences, and assuming we have MORE short sequences, for eg 20% long sequences and 80% short sequences, we get L0.2L+0.8S→L0.2L\=5\\frac{L}{0.2L + 0.8S}\\rightarrow\\frac{L}{0.2L}=50.2L+0.8SL​→0.2LL​\=5 so 5x faster training! This means packing's speedup depends on how short rows your dataset has (the more shorter, the faster). ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#padding-free-by-default) 🎬Padding-Free by Default In addition to large throughput gains available when setting `packing = True` in your `SFTConfig` , we will **automatically use padding-free batching** in order to reduce padding waste improve throughput and increases tokens/s throughput, while resulting in the _**exact same loss**_ as seen in the previous version of Unsloth. For example for Qwen3-8B and Qwen3-32B, we see memory usage decrease by 60%, be 2x faster, and have the same exact loss and grad norm curves! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FPATEJoJwIotXNPsYT1hu%252FW%2526B%2520Chart%252010_12_2025%252C%25203_57_51%2520am.png%3Falt%3Dmedia%26token%3De31ee2cd-cd6e-4fd2-9c59-7f2148179815&width=768&dpr=3&quality=100&sign=d60cf128&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FjjnXdgPSgUxL9WNzx9wc%252FW%2526B%2520Chart%252010_12_2025%252C%25203_58_19%2520am.png%3Falt%3Dmedia%26token%3D54368c73-2ce1-4faa-a1f4-c82341638be3&width=768&dpr=3&quality=100&sign=517a9289&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FA61fgCtUj0K9dhrHCt0C%252FW%2526B%2520Chart%252010_12_2025%252C%25203_54_40%2520am.png%3Falt%3Dmedia%26token%3Db8472635-4b05-430e-9df1-3820ed381c3f&width=768&dpr=3&quality=100&sign=5183f009&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FPh0Xfaup0CTz8REL6P9F%252FW%2526B%2520Chart%252010_12_2025%252C%25203_56_38%2520am.png%3Falt%3Dmedia%26token%3D86ed33c3-e5ac-4b71-82c3-f8f86ca79862&width=768&dpr=3&quality=100&sign=516b3817&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F6FKM6pkRkQX3gzQLdcIP%252FW%2526B%2520Chart%252010_12_2025%252C%25203_55_38%2520am.png%3Falt%3Dmedia%26token%3D48d2c1d3-e6f2-420c-8209-70ef247ce63d&width=768&dpr=3&quality=100&sign=8c6e34d4&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FxbhcPeELu78xh3M01xkf%252FW%2526B%2520Chart%252010_12_2025%252C%25203_56_07%2520am.png%3Falt%3Dmedia%26token%3D375c3f27-5af8-43e5-93eb-3818cb401f95&width=768&dpr=3&quality=100&sign=f7d89684&sv=2) ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#uncontaminated-packing-2-5x-faster-training) ♠️Uncontaminated Packing 2-5x faster training Real datasets can contain different sequence lengths, so increasing the batch size to 32 for example will cause padding, making training slower and use more VRAM. In the past, increasing `batch_size` to large numbers (>32) will make training SLOWER, not faster. This was due to padding - we can now eliminate this issue via `packing = True`, and so training is FASTER! When we pack multiple samples into a single one-dimensional tensor, we keep sequence length metadata around in order to properly mask samples, without leaking attention between samples. We also need the RoPE kernel described in [Fused QK RoPE Triton Kernel with packing](https://unsloth.ai/docs/blog/3x-faster-training-packing#fused-qk-rope-triton-kernel-with-packing) to allow reset position ids. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F508zS4YN2sYnYjYkt8ej%252Fimage.png%3Falt%3Dmedia%26token%3Da05f917a-f593-4abd-a834-2f3f6652ca5a&width=768&dpr=3&quality=100&sign=2db347eb&sv=2) 4 examples without packing wastes space ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8azEy9wbeF2RWSWbNdma%252Fimage.png%3Falt%3Dmedia%26token%3Da4b96567-244a-45ed-8b18-b16074bac88c&width=768&dpr=3&quality=100&sign=b07f7998&sv=2) Uncontaminated packing creates correct attention pattern By changing the ratio of 50% short sequences, and assuming we have MORE short sequences, for eg 20% long sequences and 80% long sequences, we get L0.2L+0.8S→L0.2L\=5\\frac{L}{0.2L + 0.8S}\\rightarrow\\frac{L}{0.2L}=50.2L+0.8SL​→0.2LL​\=5 so 5x faster training! This means packing's speedup depends on how short rows your dataset has (the more shorter, the faster). ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#analysis-and-benchmarks) 🏖️Analysis and Benchmarks To demonstrate the various improvements when training with our new kernels and packed data, we ran fine-tuning runs with [Qwen3-32B](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune) , Qwen3-8B, Llama 3 8B on the `yahma/alpaca-cleaned` dataset and measured various [training loss](https://unsloth.ai/docs/blog/3x-faster-training-packing#padding-free-by-default) throughput and efficiency metrics. We compared our new runs vs. a standard optimized training run with our own kernels/optimizations turned on and kernels like Flash Attention 3 (FA3) enabled. We fixed `max_length = 1024` and varied the batch size in {1, 2, 4, 8, 16, 32}. This allows the maximum token count per batch to vary in {1024, 2048, 4096, 8192, 16K, 32K}. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FFfmdok7AmeretPSGlZjg%252Fnew%2520rope%2520kernel%2520graph.png%3Falt%3Dmedia%26token%3Dd890fd95-c8c0-4817-9ee3-18e3095cde5f&width=768&dpr=3&quality=100&sign=b2e88e6b&sv=2) The above shows how tokens per second (tokens/s) training throughput varies for new Unsloth with varying batch size. This translates into training your model on an epoch of your dataset **1.7-3x faster (sometimes even 5x or more)**! These gains will be more pronounced if there are many short sequences in your data and if you have longer training runs, as described in [Why is padding needed & mathematical speedup](https://unsloth.ai/docs/blog/3x-faster-training-packing#why-is-padding-needed-and-mathematical-speedup) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FReViERxLWBHnT8GOv0ql%252Fpacking_efficiency_by_per_device_train_batch_size.png%3Falt%3Dmedia%26token%3D1c4a78c7-a611-4374-ac03-94aabb1d3184&width=768&dpr=3&quality=100&sign=2d0506c4&sv=2) The above shows the average percentage of tokens per batch that are valid (i.e., non-padding). As the batch size length grows, many more padding tokens are seen in the unpacked case, while we achieve a high packing efficiency in the packed case regardless of max sequence length. Note that, since the batching logic trims batches to the maximum sequence length seen in the batch, when the batch size is 1, the unpacked data is all valid tokens (i.e., no padding). However, as more examples are added into the batch, padding increases on average, hitting nearly 50% padding with batch size is 8! Our sample packing implementation eliminates that waste. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FRfzNZVz9uzDPhEe3frGe%252Funknown.png%3Falt%3Dmedia%26token%3De9fe893e-6b94-4c0d-b144-ef8315067c1e&width=768&dpr=3&quality=100&sign=d142afc7&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FTHtcpIdQ0z0mGRYYBfuF%252Funknown.png%3Falt%3Dmedia%26token%3Dec52b2c1-1e60-4ed2-969a-d760af26be2a&width=768&dpr=3&quality=100&sign=b85354fa&sv=2) The first graph (above) plots progress on `yahma/alpaca-cleaned` with `max_length = 2048`, Unsloth new with packing + kernels (maroon) vs. Unsloth old (gray). Both are trained with `max_steps = 500`, but we plot the x-axis in wall-clock time. Notice that we train on nearly 40% of an epoch in the packed case in the same amount of steps (and only a bit more wall-clock time) that it takes to train less than 5% of an epoch in the unpacked case. Similarly, the 2nd graph (above) plots loss from the same runs, this time plotted with training steps on the x-axis. Notice that the losses match in scale and trend, but the loss in the packing case is less variable since the model is seeing more tokens per training step. ### [](https://unsloth.ai/docs/blog/3x-faster-training-packing#how-to-enable-packing) ✨How to enable packing? **Update Unsloth first and padding free is done by default**! So all training is immediately 1.1 to 2x faster with 30% less memory usage at least and 0 change in loss curve metric! We also support Flash Attention 3 via Xformers, SDPA support, Flash Attention 2, and this works on old GPUs (Tesla T4, RTX 2080) and new GPUs like H100s, B200s etc! Sample packing works _regardless of choice of attention backend or model family_, so enjoy the same speedups previously had with these fast attention implementations! If you want to enable explicit packing, then add `packing = True` to enable up to 5x faster training! Note `packing=True` will change the training loss and will make the dataset number of rows truncated, since multiple short sequences are packed into 1 sequence. You might see the number of examples in the dataset shrink. To not get different training loss numbers, simply set `packing=False` and we will enable auto padding-free, which already makes training faster! All our notebooks are automatically faster (no need to do anything). See [Unsloth Notebooks](https://unsloth.ai/docs/get-started/unsloth-notebooks) Qwen3 14B faster: [![Logo](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2Fssl.gstatic.com%2Fcolaboratory-static%2Fcommon%2F6bc0853097d3a961eb07dc79629ca1d1%2Fimg%2Ffavicon.ico&width=20&dpr=3&quality=100&sign=b6118230&sv=2)Google Colabcolab.research.google.com](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(14B)-Reasoning-Conversational.ipynb) Llama 3.1 Conversational faster: [![Logo](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2Fssl.gstatic.com%2Fcolaboratory-static%2Fcommon%2Fcb638d8fc9309f9cbc2b5118adc2b1cf%2Fimg%2Ffavicon.ico&width=20&dpr=3&quality=100&sign=a2864477&sv=2)Google Colabcolab.research.google.com](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(1B_and_3B)-Conversational.ipynb) Thank you! If you're interested, see our [500K Context Training](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning) blog, [Memory Efficient RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl) blog and [Long Context gpt-oss](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training) blog for more topics on kernels and performance gains! [PreviousCurl & HTTP](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth) [Next500K Context Training](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning) Last updated 6 months ago Was this helpful? * [🥁Fused QK RoPE Triton Kernel with packing](https://unsloth.ai/docs/blog/3x-faster-training-packing#fused-qk-rope-triton-kernel-with-packing) * [🚃Int64 Indexing for Triton Kernels](https://unsloth.ai/docs/blog/3x-faster-training-packing#int64-indexing-for-triton-kernels) * [🧮Why is padding needed & mathematical speedup](https://unsloth.ai/docs/blog/3x-faster-training-packing#why-is-padding-needed-and-mathematical-speedup) * [🎬Padding-Free by Default](https://unsloth.ai/docs/blog/3x-faster-training-packing#padding-free-by-default) * [♠️Uncontaminated Packing 2-5x faster training](https://unsloth.ai/docs/blog/3x-faster-training-packing#uncontaminated-packing-2-5x-faster-training) * [🏖️Analysis and Benchmarks](https://unsloth.ai/docs/blog/3x-faster-training-packing#analysis-and-benchmarks) * [✨How to enable packing?](https://unsloth.ai/docs/blog/3x-faster-training-packing#how-to-enable-packing) Was this helpful? Copy Q * cos + rotate_half(Q) * sin is equivalent to Q * cos + Q @ R * sin where R is a rotation matrix [ 0, I] [-I, 0] dC/dY = dY * cos + dY @ R.T * sin where R.T is again the same [ 0, -I] but the minus is transposed. [ I, 0] Copy block_idx = tl.program_id(0) if LONG_INDEXING: offsets = block_idx.to(tl.int64) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE).to(tl.int64) n_elements = tl.cast(n_elements, tl.int64) else: offsets = block_idx * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE) Copy pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth_zoo Copy from unsloth import FastLanguageModel from trl import SFTTrainer, SFTConfig model, tokenizer = FastLanguageModel.from_pretrained( "unsloth/Qwen3-14B", ) trainer = SFTTrainer( model = model, processing_class = tokenizer, train_dataset = dataset, args = SFTConfig( per_device_train_batch_size = 1, max_length = 4096, …, packing = True, # required to enable sample packing! ), ) trainer.train() --- # How to Run Local AI Models with OpenCode | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/opencode.md) . This guide walks you through connecting **OpenCode** to [Unsloth](https://github.com/unslothai/unsloth) to run open LLMs **entirely locally.** OpenCode is an **open-source AI coding agent** that reads, modifies, and executes code across your project using a connected model. This works with any **local model** exposed through Unsloth’s **OpenAI-compatible API**, including: DeepSeek, Qwen, Gemma, and more. OpenCode acts as the client, while Unsloth loads and serves models via a local API. After setup, OpenCode connects to Unsloth, where you can select a loaded model and use it as a **coding agent**. This guide covers both setup options: * **OpenCode Desktop:** add Unsloth Studio manually as a custom OpenAI-compatible provider. * **OpenCode CLI:** launch it with `unsloth start opencode` and connect to your local model automatically. [OpenCode Setup](https://unsloth.ai/docs/integrations/opencode#installing-opencode-desktop) [Quickstart](https://unsloth.ai/docs/integrations/opencode#quickstart) In this tutorial, we’ll use `unsloth/Qwen3.6-27B-GGUF` loaded in Unsloth and access it directly inside OpenCode. Prefer a different model? Swap in any other model by loading it in Unsloth. ### [](https://unsloth.ai/docs/integrations/opencode#installing-opencode-desktop) Installing OpenCode Desktop MacOS Windows Linux #### [](https://unsloth.ai/docs/integrations/opencode#step-1-download-the-opencode-installer-for-mac) **Step 1: Download the OpenCode installer for Mac** Open `opencode.ai/download` in your browser of choice. Scroll down to the **OpenCode Desktop (Beta) ,** and click the `Download` button next to the macOS image name corresponding to your Mac's architecture (Apple Silicon or Intel). #### [](https://unsloth.ai/docs/integrations/opencode#step-2-install-opencode) Step 2: Install OpenCode Locate and double click the `OpenCode Desktop.dmg` installer file in your downloads folder. The installer window will open. Use your mouse to drag the **OpenCode** app icon on top of the **Applications** icon as shown. #### [](https://unsloth.ai/docs/integrations/opencode#step-3-launch-opencode) Step 3: Launch OpenCode Locate and double click the **OpenCode** icon under the **Applications** folder. The **OpenCode** desktop app will open and is now ready for your next action. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FZNP4G1c2hmHyBzNbGH5x%252Fopencode_interface.png%3Falt%3Dmedia%26token%3D5af679cb-504d-4299-a40e-0d8310289a90&width=768&dpr=3&quality=100&sign=6432194a&sv=2) #### [](https://unsloth.ai/docs/integrations/opencode#step-1-download-the-opencode-installer-for-windows) **Step 1: Download the OpenCode installer for Windows** Open `opencode.ai/download` in your browser. Scroll to **OpenCode Desktop (Beta)**, and click the **Download** button next to the Windows (x64) installer. #### [](https://unsloth.ai/docs/integrations/opencode#step-2-install-opencode-1) **Step 2: Install OpenCode** find the `OpenCode Desktop Installer.exe` in your Downloads folder and double click it. You can then follow the prompts in the installer window to complete the installation. #### [](https://unsloth.ai/docs/integrations/opencode#step-3-launch-opencode-1) **Step 3: Launch OpenCode** Open the Windows Start menu and search for **OpenCode**. Click the OpenCode app icon to launch it. The OpenCode desktop app will open and is now ready for your next action ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FLECAvpw6DMNCcuFclzoy%252Fimage.png%3Falt%3Dmedia%26token%3D7a628b56-3e0c-4d00-8054-ab8b9317e4f9&width=768&dpr=3&quality=100&sign=71172d9&sv=2) #### [](https://unsloth.ai/docs/integrations/opencode#step-1-download-the-opencode-installer-for-linux) **Step 1: Download the OpenCode installer for Linux** Go to `opencode.ai/download` using your preferred web browser. Find the **OpenCode Desktop (Beta)** section, then choose the Linux download option for your system. #### [](https://unsloth.ai/docs/integrations/opencode#step-2-install-opencode-2) **Step 2: Install OpenCode** Open your Downloads folder and find the **OpenCode** installer you just downloaded. For an AppImage download, right click the file, open **Properties**, and enable the option that allows the file to run as a program. Then double click the AppImage to start **OpenCode**. Or, if you prefer using the terminal, run: If you downloaded a `.deb` file, install it with: If you downloaded an `.rpm` file, install it with: #### [](https://unsloth.ai/docs/integrations/opencode#step-3-launch-opencode-2) **Step 3: Launch OpenCode** Open your Linux application launcher and search for **OpenCode**. Click the OpenCode app icon to launch it. The OpenCode desktop app will open and is now ready for your next action. ### [](https://unsloth.ai/docs/integrations/opencode#quickstart) ⚡ Quickstart After installing OpenCode, we'll need to install Unsloth Studio to enable OpenCode to serve and run inference of local models. 1. **Install or update Unsloth Studio.** Earlier versions don't expose the external API. See Installation. 2. **Launch Unsloth.** Note the port it starts on is usually `8000` or `8888`. You'll see it in the terminal output and in the browser URL (`http://localhost:PORT`). 3. **Load a model.** Click **New Chat**, pick or search a model (GGUF), and wait for it to finish loading. 4. **Connect OpenCode.** Run `unsloth start opencode`. It mints an API key, writes the config, and launches OpenCode against your loaded model. ### [](https://unsloth.ai/docs/integrations/opencode#creating-an-api-key) 🔑 Creating an API key 1. Open the sidebar, click your **Unsloth** avatar at the bottom-left. 2. Go to **Settings** → **API**. 3. Enter a friendly name (e.g. `claude-code-macbook`). 4. _(Optional)_ Set an expiry. 5. Click **Create**. 6. **Copy the key immediately.** Unsloth stores only a hash and you won't be able to view it again. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FrWnqY5PgvZdbWK8pmlD5%252Fimage.png%3Falt%3Dmedia%26token%3D3507e7d7-b552-447f-b503-2418802b8f6e&width=768&dpr=3&quality=100&sign=d4395e2c&sv=2) All keys start with the `sk-unsloth-` prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with `401 Unauthorized`. [](https://unsloth.ai/docs/integrations/opencode#connecting-unsloth-to-opencode-desktop) 🖇️ Connecting Unsloth to OpenCode Desktop ---------------------------------------------------------------------------------------------------------------------------------------- **Opencode** supports any OpenAI-compatible provider, so you can wire Unsloth in as a **Custom** provider. The setup is a one-time flow inside opencode's **Connect provider** dialog. **1\. Open the provider picker.** In opencode, type `/model` (or click the model selector at the bottom of the input). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fl59D1AVRI57fQBJqgp8t%252Fslash-model.png%3Falt%3Dmedia%26token%3D9434f804-6a9e-4b76-ab7e-75cd745fe9c2&width=768&dpr=3&quality=100&sign=4784ebab&sv=2) Then click **Connect provider** at the top-right of the select model dialog. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fi9VOd4yRAmLnqZ1D4tBm%252Fconnect%2520provider.png%3Falt%3Dmedia%26token%3Dcb28387a-9cc3-4c5f-928e-82a56edd86df&width=768&dpr=3&quality=100&sign=cec55f7&sv=2) **2\. Choose "Custom".** In the provider list, scroll to **Other** and pick **Custom**. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FWyltmg0o9YtyydUSLvcn%252Fcustom%2520highlighted.png%3Falt%3Dmedia%26token%3D5cc89686-dfae-4ebd-8beb-e6dc78088b4b&width=768&dpr=3&quality=100&sign=d0469771&sv=2) **3\. Fill in the custom provider form:** Field Value **Provider ID** `unsloth-studio` _(lowercase, hyphens allowed)_ **Display name** `Unsloth Studio` **Base URL** `http://localhost:8888/v1/` _(replace_ `_8888_` _with your_ Unsloth _port; keep the trailing_ `_/v1/_`_)_ **API key** Your `sk-unsloth-…` key In the **Models** section, add one row per model you want to expose. The left field is the model ID as Unsloth serves it; the right field is what opencode will display: Model ID (left) Display name (right) `unsloth/Qwen3.6-27B-GGUF` _(the exact name of the model as shown in Unsloth)_ `unsloth/Qwen3.6-27B-GGUF` _(shown inside opencode)_ Leave **Headers** empty unless you're proxying Unsloth through an auth layer that needs custom headers. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FjGU2OaOKRtqCwmO34BuN%252Fopencode_custom_provider_config.png%3Falt%3Dmedia%26token%3D6ef67ecf-23e6-4913-bc02-c26701f5aea0&width=768&dpr=3&quality=100&sign=edd1abb5&sv=2) **4\. Click Submit.** You should see an _"Unsloth Studio connected. Unsloth models are now available to use"_ toast. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FuCOOOJhO8brhEwwIsiUC%252Fopencode_model_available_toast.png%3Falt%3Dmedia%26token%3Da463bb0c-724a-4809-9ddb-b7bff094b255&width=768&dpr=3&quality=100&sign=abadc02f&sv=2) **Restart opencode after adding the provider.** The new provider only becomes selectable after a restart. **5\. Select your Unsloth model.** Once opencode is back up, type `/model`, search `unsloth`, and pick the model under the **Unsloth Studio** group. It'll be active on your next message. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FFRNofXnF7r7CKlHjLMwv%252Fopencode_unsloth_model_usage.png%3Falt%3Dmedia%26token%3D4c882e8f-6ac5-4f9b-8f59-5e86e312851b&width=768&dpr=3&quality=100&sign=10edad20&sv=2) Unsloth supports both OpenAI and Anthropic python SDKs. ### [](https://unsloth.ai/docs/integrations/opencode#opencode-cli) ⚙️ **OpenCode CLI** OpenCode CLI can connect to a model already running in Unsloth Studio, or start one automatically when Unsloth is not running. #### [](https://unsloth.ai/docs/integrations/opencode#if-unsloth-studio-is-already-running) If Unsloth Studio is already running With a model loaded in Unsloth Studio, open your project folder and run: Using this command launches OpenCode with the model currently loaded in Unsloth, without changing your existing OpenCode setup. #### [](https://unsloth.ai/docs/integrations/opencode#start-a-model-automatically) Start a model automatically If Unsloth Studio isn't already running, you can launch a temporary server and load a model in a single command: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fom5iWH1f42GZ69utVSyf%252Fimage.png%3Falt%3Dmedia%26token%3D815a0458-b252-4c96-a341-e345842bd248&width=768&dpr=3&quality=100&sign=4537a836&sv=2) Once **OpenCode** opens, give it a task such as: It can inspect the project, create files, edit code, and run commands using the local model. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FOA6dENYOpJdGa3zwS8Ei%252Fimage.png%3Falt%3Dmedia%26token%3D266a85f8-ec80-4a7d-bd58-e82d9ff0b864&width=768&dpr=3&quality=100&sign=c2420bd9&sv=2) In this example, the model created a two-level platformer. Let’s open it up and see how it turned out. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FHqFv9FyPsM0poo5c6WNb%252Fezgif-68bd7a9b463725d0.gif%3Falt%3Dmedia%26token%3Dabafdfef-6e13-430b-bccb-140f29ee6d13&width=768&dpr=3&quality=100&sign=9e3025ab&sv=2) #### [](https://unsloth.ai/docs/integrations/opencode#resume-your-work) Resume your work OpenCode keeps your session history automatically, so you do not need `--persist`. To pick up where you left off, even if Unsloth is no longer running, use: To open a particular session instead, use `--session `. See the complete [unsloth start](https://unsloth.ai/docs/integrations/unsloth-start) reference for model loading, remote Unsloth servers, and advanced options. ### [](https://unsloth.ai/docs/integrations/opencode#optional-configure-server-access) Optional: configure server access `unsloth run` starts the local API server and loads a model for OpenCode to connect to. You can also customize how the server behaves when starting it. Use `--disable-tools` when driving OpenCode (or any external coding agent). By default Unsloth Studio runs its own server-side tools, which swallows the agent's tool calls, so OpenCode answers but never edits files. `--disable-tools` switches to passthrough, so OpenCode's own tools are used. Use `-p` to change which port the server runs on. This starts the server on `0.0.0.0:8888`, allowing other devices on your local network to connect. For more advanced runtime configuration, see the main [API tuning](https://unsloth.ai/docs/basics/api#unsloth-run-command) section. [PreviousOpenClaw](https://unsloth.ai/docs/integrations/openclaw) [NextPython SDK](https://unsloth.ai/docs/integrations/connect-python-sdk-to-unsloth) Last updated 3 days ago Was this helpful? * [Installing OpenCode Desktop](https://unsloth.ai/docs/integrations/opencode#installing-opencode-desktop) * [⚡ Quickstart](https://unsloth.ai/docs/integrations/opencode#quickstart) * [🔑 Creating an API key](https://unsloth.ai/docs/integrations/opencode#creating-an-api-key) * [🖇️ Connecting Unsloth to OpenCode Desktop](https://unsloth.ai/docs/integrations/opencode#connecting-unsloth-to-opencode-desktop) * [⚙️ OpenCode CLI](https://unsloth.ai/docs/integrations/opencode#opencode-cli) * [Optional: configure server access](https://unsloth.ai/docs/integrations/opencode#optional-configure-server-access) Was this helpful? Copy chmod +x OpenCode*.AppImage ./OpenCode*.AppImage Copy sudo apt install ./OpenCode*.deb Copy sudo dnf install ./OpenCode*.rpm Copy unsloth start opencode Copy unsloth start opencode --model unsloth/gemma-4-26B-A4B-it-GGUF Copy Create a Flash-style Python 2D game with jumping, dashing, and two levels. Copy unsloth start opencode \ --model unsloth/Qwen3.6-27B-GGUF \ --continue Copy # Run the API on port 8888 (--disable-tools passes OpenCode's own tools through) unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ --disable-tools \ -p 8888 Copy # Allow other devices on your network to connect unsloth run \ --model unsloth/gemma-4-26B-A4B-it-GGUF \ -H 0.0.0.0 \ --disable-tools \ -p 8888 --- # How to Connect Ollama to Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connections/ollama.md) . Ollama lets you run local LLMs on your own hardware, and [Unsloth](https://github.com/unslothai/unsloth) makes it easy to connect and run those models directly into a open-source UI chat interface. In this guide, you’ll learn how to install Ollama, run native Ollama models or GGUF models from Hugging Face, connect Ollama to Unsloth, and start chatting with local AI models. Whether you want to use models like [Qwen](https://unsloth.ai/docs/models/qwen3.6) , import a GGUF file, or expose your local Ollama server through an OpenAI-compatible endpoint, this walkthrough covers the full setup from installation to first chat. ### [](https://unsloth.ai/docs/integrations/connections/ollama#setup) Setup 1 #### [](https://unsloth.ai/docs/integrations/connections/ollama#install-or-prepare-ollama) Install or prepare Ollama macOS Windows Linux Docker Install Ollama with the install script: Copy curl -fsSL https://ollama.com/install.sh | sh You can also download Ollama manually from [ollama.com/download](https://ollama.com/download) . Install Ollama from PowerShell: Copy irm https://ollama.com/install.ps1 | iex You can also download Ollama manually from [ollama.com/download](https://ollama.com/download/OllamaSetup.exe) . Install Ollama with the install script: Copy curl -fsSL https://ollama.com/install.sh | sh You can also download Ollama manually from [ollama.com/download](https://docs.ollama.com/linux#manual-install) . The official Ollama Docker image is `ollama/ollama` on Docker Hub. Copy docker run -d \ -v ollama:/root/.ollama \ -p 11434:11434 \ --name ollama \ ollama/ollama Ollama usually runs at: Copy http://localhost:11434 2 #### [](https://unsloth.ai/docs/integrations/connections/ollama#run-a-model) Run a model You can choose a model in two common ways: * Search native Ollama models at [ollama.com/search](https://ollama.com/search) , then copy the model name. * Use a GGUF model from Hugging Face, then copy the Ollama command from **Use this model**. For an Ollama model, pull and run it: Copy ollama pull qwen3.6:35b-a3b ollama run qwen3.6:35b-a3b If the Ollama app or service is not already running, start it first: Copy ollama serve #### [](https://unsloth.ai/docs/integrations/connections/ollama#pick-a-gguf-from-hugging-face) Pick a GGUF from Hugging Face If you are using a GGUF model from Hugging Face, the easiest way to get the command is from the model page. Open the model you want to use, click **Use this model**, then choose **Ollama** from the local apps list. Pick the quantization you want from the dropdown, then copy the generated command. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FRSR0qWpzSRJgeSIghoWO%252Fexport-1779051800964-small.gif%3Falt%3Dmedia%26token%3D8e557e60-2e11-420e-a7f4-dd80533bd86b&width=768&dpr=3&quality=100&sign=d15c347e&sv=2) For example, with Ollama: Copy ollama run hf.co/unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL This helps avoid mistakes with the repo name or quantization tag. 3 #### [](https://unsloth.ai/docs/integrations/connections/ollama#connect-ollama-to-unsloth) Connect Ollama to Unsloth Open **Settings → Connections**, then click **Add Connection**. Select **Ollama**, then enter your connection details: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FIKDXfJZmRQVA742rXAHA%252Fimage.png%3Falt%3Dmedia%26token%3D6c7105a0-9f6e-435b-bafb-1502e22331bd&width=768&dpr=3&quality=100&sign=90caa5c4&sv=2) Use the Ollama URL shown in the Unsloth form. In most local setups, this is: Copy http://localhost:11434 If Unsloth asks for an OpenAI-compatible base URL, use: Copy http://localhost:11434/v1 Ollama normally does not need an API key. Leave the API key field empty unless you are using a proxy that requires one. Click **Load Models** to fetch the models running in Ollama, or enter the **model ID** yourself, for example `qwen3.6`. 4 #### [](https://unsloth.ai/docs/integrations/connections/ollama#ready-to-chat) Ready to Chat After you click **Add Connection**, the models you enabled will now appear under **Connected** in the **Select Model** dropdown. #### [](https://unsloth.ai/docs/integrations/connections/ollama#common-ollama-commands) Common Ollama commands Use these while setting up the model you want to expose to Unsloth: Command What it does `ollama run qwen3.6:35b-a3b` Run a model and open an interactive chat `ollama pull qwen3.6:35b-a3b` Download a model without starting chat `ollama ls` List downloaded models `ollama ps` List models currently running `ollama stop qwen3.6:35b-a3b` Stop a running model `ollama rm qwen3.6:35b-a3b` Remove a downloaded model `ollama serve` Start the Ollama server If you are importing a local GGUF into Ollama, create a `Modelfile`, then run: Copy ollama create -f Modelfile If Ollama is not detected, make sure the Ollama app or service is running. Then click **Load Models** again in Unsloth. For the full command list, see the [Ollama CLI reference](https://docs.ollama.com/cli) . [PreviousvLLM](https://unsloth.ai/docs/integrations/connections/vllm) [NextOpenRouter](https://unsloth.ai/docs/integrations/connections/openrouter) Last updated 2 months ago Was this helpful? Was this helpful? --- # Run Coding Agents with Local LLMs using Unsloth Start | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/unsloth-start.md) . Unsloth lets you connect [Claude Code](https://unsloth.ai/docs/basics/claude-code) , [Codex](https://unsloth.ai/docs/basics/codex) , Hermes, OpenCode, Pi, and other coding agents to a local model via the `unsloth start` command. The entire workflow can run offline on your own hardware. [Unsloth Studio](https://unsloth.ai/docs/new/studio) automatically configures the endpoint, API key, provider, model, and context length for each launch, so you can use your preferred agent without modifying them. This guide will show you how to launch models from the command line for offline use, and connect to Unsloth Studio. ### [](https://unsloth.ai/docs/integrations/unsloth-start#quickstart) Quickstart First, make sure you have [Unsloth installed](https://unsloth.ai/docs/new/studio/install) . Then open Unsloth, load a model, go to your project folder, and run the command in terminal: Copy unsloth start claude You can replace `claude` with any agent below: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpXE6kCHjh8qOEaggf94M%252FScreenshot_20260718_122426.png%3Falt%3Dmedia%26token%3Da59e4c8c-efdb-451b-b1f8-621955564f6d&width=768&dpr=3&quality=100&sign=4c664a42&sv=2) Claude Code running with Qwen3.5 locally. Agent Command Claude Code `unsloth start claude` OpenAI Codex `unsloth start codex` Hermes Agent `unsloth start hermes` OpenClaw `unsloth start openclaw` OpenCode `unsloth start opencode` Pi Coding Agent `unsloth start pi` Unsloth uses temporary or session-scoped provider configuration. It does not add an Unsloth provider to the agent's normal configuration files. Codex currently requires a GGUF model served through the `llama-server` backend. ### [](https://unsloth.ai/docs/integrations/unsloth-start#agent-guides) Agent Guides [](https://unsloth.ai/docs/basics/claude-code) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FT8nv3UJzIrcxmjSn8OfW%252Fclaude-code.webp%3Falt%3Dmedia%26token%3D811f9635-7321-427c-a11e-29a40e49e376&width=490&dpr=3&quality=100&sign=bfcb8be5&sv=2) Claude Code [](https://unsloth.ai/docs/basics/codex) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FLxRuVHEBYywquhRGBYu0%252Fcodex%2520only%2520logo.png%3Falt%3Dmedia%26token%3D7faa9cbd-090a-45dc-9d5e-acb1a9591f27&width=490&dpr=3&quality=100&sign=b45fab1e&sv=2) Codex [](https://unsloth.ai/docs/integrations/hermes-agent) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F6ysPTzsXIO5diu4vOTCh%252Fimages.jpeg%3Falt%3Dmedia%26token%3D5cc58141-69eb-4450-8214-7a8421fbe2a5&width=490&dpr=3&quality=100&sign=61ffa1a4&sv=2) Hermes Agent [](https://unsloth.ai/docs/integrations/openclaw) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FAd5uo6LjavtBrrKZfwZ3%252Fopenclaw-hero-light.png%3Falt%3Dmedia%26token%3D3097e40f-e80c-4c9c-8efc-02ecf5530ef6&width=490&dpr=3&quality=100&sign=19c2bd99&sv=2) OpenClaw [](https://unsloth.ai/docs/integrations/opencode) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FlCi6Ab4f2qLlNj2EWGL0%252Fopencodelogo.png%3Falt%3Dmedia%26token%3D6118773e-5392-4718-a431-a50157a34cbe&width=490&dpr=3&quality=100&sign=24ebad1b&sv=2) OpenCode [](https://unsloth.ai/docs/basics/api) ![Cover](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FoRer1Vbuh9wSPTawz0ey%252Funsloth%2520api%2520long%2520logo.png%3Falt%3Dmedia%26token%3D860a8143-9d59-466b-abed-6ea7d67e51fd&width=490&dpr=3&quality=100&sign=5b60331c&sv=2) Unsloth API ### [](https://unsloth.ai/docs/integrations/unsloth-start#load-a-model-from-the-command-line) Load a model from the command line ![unsloth start launching Codex with a local GGUF model in Unsloth Studio](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FAOqECGHODYPDGeYDpXOz%252Fwasim%2520unslothstart.png%3Falt%3Dmedia%26token%3Da7a1bf5e-d964-4bd4-a0d0-f86753e9f5a8&width=768&dpr=3&quality=100&sign=8c27b73b&sv=2) Unsloth start finds or loads the model, configures the coding agent and launches it from the current project. You can select and load a model while launching the agent: With quant suffix With \`--gguf-variant\` The `:UD-Q4_K_XL` suffix selects the GGUF quant. An explicit `--gguf-variant` overrides the quant written after the model name. On the default local address, passing `--model` lets `unsloth start` start a temporary server when Unsloth Studio is not already running. The temporary server stops when the agent exits. If Unsloth is already running, the command connects to it and leaves it running. See below for examples of agents being connected to a local LLM: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fmhay2zcaqL5mNNBeYwKv%252FScreenshot_20260717_163314.png%3Falt%3Dmedia%26token%3D5a1e2f04-92b1-42f1-adac-9494fb43613a&width=768&dpr=3&quality=100&sign=1604c9e6&sv=2) OpenCode ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FnrT2lJtpD4uwMkn5HAhO%252FScreenshot_20260717_163427.png%3Falt%3Dmedia%26token%3Dffe77a32-05a9-439e-845f-5a7a4b034464&width=768&dpr=3&quality=100&sign=7d71d6c9&sv=2) Hermes ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FfKxk4eUVf0tZB1IFtava%252FScreenshot_20260717_161638.png%3Falt%3Dmedia%26token%3Db709f099-704e-4d97-947e-d75cbf43cd8a&width=768&dpr=3&quality=100&sign=1a927308&sv=2) Claude Code ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FxPckgHYLacJK1jpwXHOR%252FScreenshot_20260717_163145.png%3Falt%3Dmedia%26token%3De61f7404-341e-474b-bc08-59adc66ce274&width=768&dpr=3&quality=100&sign=2e522d57&sv=2) Codex ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fy5QNC3RaSkCIr1KFBrMd%252FScreenshot_20260717_163831.png%3Falt%3Dmedia%26token%3De1501dbe-6a2b-4a4d-96df-92a3e8f2f3df&width=768&dpr=3&quality=100&sign=134779f2&sv=2) Openclaw ### [](https://unsloth.ai/docs/integrations/unsloth-start#connect-to-a-remote-unsloth-server) Connect to a remote Unsloth server Set the Unsloth URL and API key before launching an agent: You can also pass the key with `--api-key`. For a verified local Unsloth server, `unsloth start` creates or reuses the API key automatically. ### [](https://unsloth.ai/docs/integrations/unsloth-start#options) Options Option What it does `--model`, `-m` Select a model. Without it, use the first model reported by Unsloth. `--api-key` Supply a Unsloth API key. You can also set `UNSLOTH_API_KEY`. `--launch` / `--no-launch` Launch the agent or print the generated environment and command. `--serve` / `--no-serve` Allow or prevent automatic local server startup. `--gguf-variant` Select a GGUF quantization variant. `--context-length`, `--max-seq-length` Set the requested context length when loading the model. `--load-in-4bit` / `--no-load-in-4bit` Control 4-bit loading for non-GGUF Hugging Face models. `--tensor-parallel` / `--no-tensor-parallel` Enable or disable tensor-parallel GGUF loading for multi-GPU systems. `--persist` / `--no-persist` Keep the Unsloth-managed agent home where applicable. `--yolo` Use the selected agent's non-prompting or trust mode. `-h`, `--help` Show command help. Load options are used when Unsloth needs to load or reconcile the requested model. ### [](https://unsloth.ai/docs/integrations/unsloth-start#pass-normal-commands-to-the-agent) Pass normal commands to the agent Arguments that are not Unsloth options are passed to the selected agent: Use the agent's own help command for its complete list of native options. ### [](https://unsloth.ai/docs/integrations/unsloth-start#sessions-and-persist) Sessions and `--persist` `--persist` keeps managed storage. It does not resume a conversation by itself; also pass the agent's normal resume command. Agent Do you need `--persist`? Resume example Claude Code No. Claude uses its normal session store. `unsloth start claude --continue` OpenAI Codex Yes. Its managed Codex home is temporary by default. `unsloth start codex --persist resume --last` OpenClaw Yes. Managed config, workspace and sessions are temporary by default. Use `--persist` with the same native session ID. OpenCode No. OpenCode uses its normal session store. `unsloth start opencode run --continue "Continue"` Hermes Agent Yes. Its managed Hermes home is temporary by default. `unsloth start hermes --persist --continue --oneshot "Continue"` Pi Coding Agent Yes. Its managed Pi home is temporary by default. `unsloth start pi --persist --continue` For Codex, OpenClaw, Hermes and Pi, use `--persist` from the first launch and again when returning to the session. Codex OpenClaw Hermes Agent Pi Coding Agent Use a stable session ID for a named session: #### [](https://unsloth.ai/docs/integrations/unsloth-start#print-the-launch-command-without-running-it) Print the launch command without running it This prints the generated environment and command. The output can contain connection credentials, so do not publish it in logs or screenshots. ### [](https://unsloth.ai/docs/integrations/unsloth-start#permission-bypass-mode) Permission-bypass mode `--yolo` maps to the selected agent's trust or non-prompting mode. It can reduce approval prompts and allow the agent to run commands without asking first. Only use it in an environment where unrestricted agent actions are acceptable. #### [](https://unsloth.ai/docs/integrations/unsloth-start#common-issues) Common issues No Unsloth server was found[](https://unsloth.ai/docs/integrations/unsloth-start#no-unsloth-server-was-found) Start Unsloth and load a model, or pass \`--model\` so Unsloth can start a temporary local server. Codex does not connect[](https://unsloth.ai/docs/integrations/unsloth-start#codex-does-not-connect) Use a GGUF model running through \`llama-server\`. The current Codex integration does not use the Unsloth transformers backend. A session was not restored[](https://unsloth.ai/docs/integrations/unsloth-start#a-session-was-not-restored) For Codex, OpenClaw, Hermes and Pi, use \`--persist\` on the first and later launches, then pass the agent's native resume option. [PreviousContinued Pretraining](https://unsloth.ai/docs/basics/continued-pretraining) [NextConnect a Provider](https://unsloth.ai/docs/integrations/connections) Last updated 1 day ago Was this helpful? * [Quickstart](https://unsloth.ai/docs/integrations/unsloth-start#quickstart) * [Agent Guides](https://unsloth.ai/docs/integrations/unsloth-start#agent-guides) * [Load a model from the command line](https://unsloth.ai/docs/integrations/unsloth-start#load-a-model-from-the-command-line) * [Connect to a remote Unsloth server](https://unsloth.ai/docs/integrations/unsloth-start#connect-to-a-remote-unsloth-server) * [Options](https://unsloth.ai/docs/integrations/unsloth-start#options) * [Pass normal commands to the agent](https://unsloth.ai/docs/integrations/unsloth-start#pass-normal-commands-to-the-agent) * [Sessions and --persist](https://unsloth.ai/docs/integrations/unsloth-start#sessions-and-persist) * [Permission-bypass mode](https://unsloth.ai/docs/integrations/unsloth-start#permission-bypass-mode) Was this helpful? Copy unsloth start codex \ --model unsloth/gemma-4-E2B-it-GGUF:UD-Q4_K_XL \ --context-length 32768 Copy unsloth start codex \ --model unsloth/gemma-4-E2B-it-GGUF \ --gguf-variant UD-Q4_K_XL \ --context-length 32768 Copy export UNSLOTH_STUDIO_URL=https://studio.example.com export UNSLOTH_API_KEY=sk-unsloth-... unsloth start claude Copy unsloth start claude --continue unsloth start codex --persist resume --last unsloth start opencode run --continue "Continue the previous task" unsloth start pi --persist --continue Copy unsloth start codex --persist unsloth start codex --persist resume --last Copy unsloth start openclaw --persist \ agent --local --session-id my-session --message "Inspect this repository" unsloth start openclaw --persist \ agent --local --session-id my-session --message "Continue" Copy unsloth start hermes --persist --oneshot "Inspect this repository" unsloth start hermes --persist --continue --oneshot "Continue" Copy unsloth start pi --persist --print "Inspect this repository" unsloth start pi --persist --continue --print "Continue" Copy unsloth start claude --no-launch --- # Connect Curl & HTTP to Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth.md) . Unsloth exposes three OpenAI/Anthropic-compatible wire formats at the same base URL on the port Unsloth started on. All of them take an `Authorization: Bearer sk-unsloth-…` header and return either JSON or SSE, depending on whether you set `stream`. This page groups the recipes by endpoint (`/v1/chat/completions`, `/v1/messages`, `/v1/responses`, `/v1/models`) and ends with a shared section on Unsloth's built-in **server-side tools**, which work across all the chat endpoints. If you're not sure what URL / key / model name to use, read the API overview first. It walks you through starting Unsloth, loading a model, and creating an `sk-unsloth-…` key. ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#authentication) 🔑 Authentication Every request needs an `Authorization` header: Copy Authorization: Bearer sk-unsloth-xxxxxxxxxxxx To keep keys out of your shell history, export the key once and reference the env var: Copy export UNSLOTH_STUDIO_AUTH_TOKEN=sk-unsloth-xxxxxxxxxxxx The snippets below inline the key as `sk-unsloth-xxxxxxxxxxxx` for clarity. In practice, substitute `$UNSLOTH_STUDIO_AUTH_TOKEN`. ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#list-loaded-models) 📋 List loaded models Copy curl http://localhost:8888/v1/models \ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" Response: Copy { "object": "list", "data": [\ {"id": "unsloth/gemma-3-27b-it-GGUF", "object": "model", "owned_by": "local"}\ ] } ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FgVC1VPkCkUQBVg0HScEc%252FScreenshot%25202026-04-22%2520at%25203.06.57%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D6f77cc78-bd0e-4141-b555-b232df1036dc&width=768&dpr=3&quality=100&sign=e07f20f1&sv=2) Use the `id` field whenever a request needs a `"model"` value (or when a client like opencode asks for a **Model ID**). ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#chat-completions-v1-chat-completions) 💬 Chat Completions (`/v1/chat/completions`) The OpenAI Chat Completions dialect. The broadest compatibility surface. Works with the OpenAI SDK, opencode, Cursor, Continue, Cline, Open WebUI, SillyTavern, and most OpenAI-compatible tools. #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#basic-request) Basic request ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fvy9TyeJZljKrMJAath0L%252FScreenshot%25202026-04-22%2520at%25203.40.35%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D98f8d622-83c0-4f55-aa37-d3571f0a38cd&width=768&dpr=3&quality=100&sign=b357535f&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#streaming) Streaming Add `"stream": true` and the response switches to Server-Sent Events (`text/event-stream`). Tell `curl` to flush as bytes arrive with `--no-buffer` (`-N`): Each line of the response looks like `data: {"choices":[{"delta":{"content":"..."}}]}`, ending with `data: [DONE]`. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FsazSu4RpyiEa6HeamYbj%252FScreenshot%25202026-04-22%2520at%25203.42.05%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Df62cf129-74cd-4e8e-986a-1755eeb0e048&width=768&dpr=3&quality=100&sign=a01b5fba&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#images-vision) Images (vision) Attach an image as an `image_url` content part in the user message. The URL can be HTTPS or a base64 `data:` URI: The loaded model must be multimodal. If you load a text-only model the request succeeds structurally but the model won't process the image. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F9yO7ersnyaIj3X8UkD5q%252FScreenshot%25202026-04-22%2520at%25203.59.30%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D9e0af9e5-d565-4acd-8b4a-020c02d74cd8&width=768&dpr=3&quality=100&sign=583769b4&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#function-calling-openai-tools) Function calling (OpenAI tools) Pass OpenAI-style `tools` and (optionally) `tool_choice`. Your client runs each tool call and returns the result on the next turn. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F39CDJk0Z1WhG5JyaJIfV%252Ftests_codex_pr_2.png%3Falt%3Dmedia%26token%3Dc3f01286-7dc9-47dd-b7f9-9ee8c2d3cb2f&width=768&dpr=3&quality=100&sign=23f73c44&sv=2) ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#anthropic-messages-v1-messages) 📨 Anthropic Messages (`/v1/messages`) Unsloth's Anthropic-compatible dialect used by Claude Code, the Anthropic SDK, OpenClaw, and any client that speaks the Messages API. #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#basic-request-1) Basic request `max_tokens` is required on `/v1/messages` (it's optional on `/v1/chat/completions`). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FOz7xeC4dRnDjO5GDAsmC%252FScreenshot%25202026-04-22%2520at%25204.01.52%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Db7025ce1-e3a8-46ad-91fa-e87b96a22d6b&width=768&dpr=3&quality=100&sign=8f45151b&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#streaming-1) Streaming Events follow Anthropic's SSE shape: `message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta`, `message_stop`, plus Unsloth's custom `tool_result` event for server-side tool output. #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#images-vision-1) Images (vision) Anthropic-style image content uses a `source` block with base64 data: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FMkzYMtkw9dt0V7cbqduU%252FScreenshot%25202026-04-22%2520at%25204.03.28%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dde854c5a-b399-4736-83bb-23ffe26981e2&width=768&dpr=3&quality=100&sign=231034a5&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#tool-calling-anthropic-tools) Tool calling (Anthropic tools) `tool_choice` values map as follows to the OpenAI dialect: Anthropic `auto` → OpenAI `auto`, Anthropic `any` → OpenAI `required`, Anthropic `{type: "tool", name: "x"}` → OpenAI `{type: "function", function: {name: "x"}}`, Anthropic `none` → OpenAI `none`. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fqb37JN1ffgBFBKyd1uWD%252FScreenshot%25202026-04-22%2520at%25204.07.43%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D1fb7ec5d-8326-4e9f-b752-2d5d7e830a39&width=768&dpr=3&quality=100&sign=c703041b&sv=2) ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#responses-v1-responses) 🧬 Responses (`/v1/responses`) Unsloth also speaks the newer **OpenAI Responses API**, the protocol Codex and other recent OpenAI clients have moved to. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FhmyMSCTgdD1NYiLUdYNZ%252Fopenai_responses.png%3Falt%3Dmedia%26token%3Ddd56bd52-71db-46f3-8f6c-1499df8e443c&width=768&dpr=3&quality=100&sign=4da7a6df&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FZ3MEH4vV2RjxhLe0DsWJ%252Fresponses_function_calling.png%3Falt%3Dmedia%26token%3Da8637d5a-51d2-47f5-9658-add9cdb757e3&width=768&dpr=3&quality=100&sign=b3da3bb2&sv=2) Streaming works the same way as Chat Completions. Add `"stream": true` and pipe with `-N`. ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#unsloth-server-side-tools-shorthand) 🧰 Unsloth server-side tools (shorthand) In addition to client-side function calling, Unsloth can execute **Python**, **bash**, and **web search** server-side and stream the results back as custom `tool_result` events. This is the feature that makes Unsloth feel like a "real" agent out of the box, no round-tripping tool calls through your client. Opt in by passing these extra fields to **either** `/v1/chat/completions` or `/v1/messages`: Field Type Notes `enable_thinking` `boolean` `false` to disable thinking. `true` by default `enable_tools` `boolean` `true` to enable server-side tool execution. `enabled_tools` `array` Which tools the model can call. Supports `python`, `bash`, `web_search`. `session_id` `string` Optional. Persists tool state (e.g. Python kernel) across calls. #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#thinking-mode) Thinking mode Thinking mode is enabled by default. The model will think before providing an answer. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F4XHr9KJhzpFEibCGEuHF%252FScreenshot%25202026-04-22%2520at%252012.22.46%25E2%2580%25AFPM.png%3Falt%3Dmedia%26token%3D15fa68ae-0ca5-464f-af16-25cfabc84c33&width=768&dpr=3&quality=100&sign=7f034a19&sv=2) To disable thinking pass `enable_thinking: false` in your request. The model will provide an answer without thinking first. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FKQR2nh8sb5G1zUc7YSs9%252FScreenshot%25202026-04-22%2520at%252012.25.47%25E2%2580%25AFPM.png%3Falt%3Dmedia%26token%3D4646a17e-2217-4f00-9e39-5362f989126f&width=768&dpr=3&quality=100&sign=af3c1b28&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#python-execution) Python execution ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FVGbYtwBfSoqgZov5ivMA%252FScreenshot%25202026-04-22%2520at%25204.16.13%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D28b83469-24fb-412d-9489-a0c6b92dd8b5&width=768&dpr=3&quality=100&sign=caa66db6&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#web-search--python-streaming) Web search + Python (streaming) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FzSFyzwXtJ0Hgvqq7gZcV%252FScreenshot%25202026-04-22%2520at%25204.25.13%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Da8fc5908-cb11-4ec9-85c6-77a6b5c2e205&width=768&dpr=3&quality=100&sign=cb62452b&sv=2) #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#on-v1-messages) On `/v1/messages` The same shorthand works against the Anthropic Messages endpoint: ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FPpJxE60f2J59T5o85tQ1%252FScreenshot%25202026-04-22%2520at%25204.28.43%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D870d74cb-dd8e-42c1-adeb-287807b22e16&width=768&dpr=3&quality=100&sign=d6a5349&sv=2) Unsloth streams its own `tool_result` SSE events in addition to the standard Anthropic / OpenAI event types, The model sees each tool's output on its next turn. ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#troubleshooting) ❔ Troubleshooting `**401 Unauthorized**` **\-** The `Authorization` header is missing or the key is wrong. Double-check: `Authorization: Bearer sk-unsloth-…`. `**curl**` **hangs on streaming requests -** Add `-N` (same as `--no-buffer`). Without it, `curl` buffers the SSE stream and you see nothing until the end. **Base64 encoding differs between OSes** **\-** Linux's `base64` defaults to wrapping lines, macOS / BSD does not. Use `base64 -w 0` on Linux, `base64` on macOS, or pipe the output through `tr -d '\n'`. **JSON escaping in shells** **\-** Heredocs (`-d @file.json`) are cleaner than inline strings once the body gets complex. Example: `curl ... -d @body.json`. `**max_tokens**` **errors on** `**/v1/messages**` **\-** The Anthropic dialect requires it. Add `"max_tokens": 1024` (or whatever limit you want). For endpoint-level issues (model not loading, connection dropped, wrong port) see the API overview page. ### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#optional-tune-server-defaults) Optional: tune server defaults You can customize default behavior when starting the server with `unsloth run`. Use `--reasoning off` to turn thinking off, or `--reasoning on` to turn it on for models that support reasoning. This starts the server on `0.0.0.0:8888`, allowing other devices on your local network to connect. #### [](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#override-settings-per-request) Override settings per request You can also override generation settings directly in each API request. Request-level values like `temperature`, `top_p`, `max_tokens`, and `stream` override the server defaults for that request. [PreviousPython SDK](https://unsloth.ai/docs/integrations/connect-python-sdk-to-unsloth) [NextNew 3x Faster Training](https://unsloth.ai/docs/blog/3x-faster-training-packing) Last updated 1 month ago Was this helpful? * [🔑 Authentication](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#authentication) * [📋 List loaded models](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#list-loaded-models) * [💬 Chat Completions (/v1/chat/completions)](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#chat-completions-v1-chat-completions) * [📨 Anthropic Messages (/v1/messages)](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#anthropic-messages-v1-messages) * [🧬 Responses (/v1/responses)](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#responses-v1-responses) * [🧰 Unsloth server-side tools (shorthand)](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#unsloth-server-side-tools-shorthand) * [❔ Troubleshooting](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#troubleshooting) * [Optional: tune server defaults](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth#optional-tune-server-defaults) Was this helpful? Copy curl http://localhost:8888/v1/chat/completions \ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \ -H "Content-Type: application/json" \ -d '{ "model": "default", "messages": [{"role": "user", "content": "Hello"}] }' Copy curl -N http://localhost:8888/v1/chat/completions \ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen-local", "messages": [{"role": "user", "content": "Write a haiku about locally-run LLMs."}], "stream": true }' Copy # Embed a local file as base64 (trimmed for brevity) IMG=$(base64 -w 0 test.jpg) cat > /tmp/request.json < /tmp/request.json < For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/mac.md). # Install Unsloth on MacOS To install Unsloth locally on your local Apple MacOS device, follow the steps below: ### Install Unsloth \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Use the same command to \*\*update\*\*. ### Launch Every time you want to launch Unsloth again: \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` For detailed Unsloth Studio install instructions and requirements, \[view our guide\](/docs/new/studio/install.md). ### Uninstall The recommended way to fully remove Unsloth Studio is the uninstall script for your OS. It stops any running servers, removes the app, CLI command, launcher data, shortcuts, and platform-specific entries (macOS \`.app\` bundle + Launch Services; Windows Start Menu + registry + PATH): \`\`\`shellscript curl -fsSL https://raw.githubusercontent.com/unslothai/unsloth/main/scripts/uninstall.sh | sh \`\`\` #### Manual uninstall If you prefer to remove only specific parts: \*\*1. Remove app only\*\* (keeps history, chats, checkpoints, and exports intact): \* \`rm -rf ~/.unsloth/studio/unsloth\_studio\` \*\*2. Remove Unsloth entirely\*\* (keeps other Unsloth tools intact): \* \`rm -rf ~/.unsloth/studio\` \*\*3. Remove everything Unsloth-related:\*\* \* \`rm -rf ~/.unsloth\` {% hint style="warning" %} Note: Step 3 deletes everything in history, chats, model checkpoints, and exports. This cannot be undone. {% endhint %} \*\*4. Remove shortcuts and symlinks:\*\* \`\`\`shellscript rm -rf ~/Applications/Unsloth\\ Studio.app ~/Desktop/Unsloth\\ Studio \`\`\` \*\*5. Remove the CLI command:\*\* \* \`rm -f ~/.local/bin/unsloth\` {% hint style="info" %} Note: Steps 1-5 dont touch your downloaded HF model files. See Deleting cached HF model files below if you want to reclaim that space. {% endhint %} ### \*\*Deleting model files\*\* You can delete old model files either from the bin icon in model search or by removing the relevant cached model folder from the Hugging Face cache directory. The default cache location is: \`\`\`bash ~/.cache/huggingface/hub/ \`\`\` If \`HF\_HUB\_CACHE\` or \`HF\_HOME\` is set, use that location instead. You can check with: \`\`\`bash echo ${HF\_HUB\_CACHE:-${HF\_HOME:-${XDG\_CACHE\_HOME:-$HOME/.cache}/huggingface}/hub} \`\`\` To delete a specific model, remove its folder (e.g. \`models--unsloth--Llama-3.1-8B-bnb-4bit\`) from the cache directory. To clear all cached models: \`\`\`bash rm -rf ~/.cache/huggingface/hub/ \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/mac.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/docker.md). # Install Unsloth via Docker Learn how to use our Docker containers with all dependencies pre-installed for immediate installation. No setup required, just run and start training! Unsloth Docker image: \[\*\*\`unsloth/unsloth\`\*\*\](https://hub.docker.com/r/unsloth/unsloth) {% hint style="success" %} Unsloth Studio now shares the same cache as notebooks and scripts to avoid unnecessary re-downloads. {% endhint %} ### ⚡ Quickstart {% stepper %} {% step %} \*\*Install Docker and NVIDIA Container Toolkit.\*\* Install Docker via \[Linux\](https://docs.docker.com/engine/install/) or \[Desktop\](https://docs.docker.com/desktop/) (other).\\ Then install \[NVIDIA Container Toolkit\](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#installation): export NVIDIA_CONTAINER_TOOLKIT_VERSION=1.17.8-1 sudo apt-get update && sudo apt-get install -y \ nvidia-container-toolkit=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ nvidia-container-toolkit-base=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ libnvidia-container-tools=${NVIDIA_CONTAINER_TOOLKIT_VERSION} \ libnvidia-container1=${NVIDIA_CONTAINER_TOOLKIT_VERSION} ![](https://unsloth.ai/files/RboTfJ1xSzZTCqQtIMLU) {% endstep %} {% step %} \*\*Run the container.\*\* \[\*\*\`unsloth/unsloth\`\*\*\](https://hub.docker.com/r/unsloth/unsloth) is Unsloth's only Docker image. \`\`\`bash docker run -d -e JUPYTER\_PASSWORD="mypassword" \\ -p 8888:8888 -p 8000:8000 -p 2222:22 \\ -v $(pwd)/work:/workspace/work \\ --gpus all \\ unsloth/unsloth \`\`\` ![](https://unsloth.ai/files/mubtsa9Nk3VMiHxCZgDK) {% endstep %} {% step %} \*\*Access Jupyter Lab\*\* Go to \[http://localhost:8888\](http://localhost:8888/) and open Unsloth. ![](https://unsloth.ai/files/znEiZdGt2Ugx2t3e9Hlx) Access the \`unsloth-notebooks\` tabs to see Unsloth notebooks. ![](https://unsloth.ai/files/DiHKL6V9wzIV1Z3MyBL8) ![](https://unsloth.ai/files/igv6nguflqvQmwHcT9Xe) {% endstep %} {% step %} \*\*Start training with Unsloth\*\* If you're new, follow our step-by-step \[Fine-tuning Guide\](/docs/get-started/fine-tuning-llms-guide.md), \[RL Guide\](/docs/get-started/reinforcement-learning-rl-guide.md) or just save/copy any of our premade \[notebooks\](/docs/get-started/unsloth-notebooks.md). ![](https://unsloth.ai/files/3IlCi9DRXN3i3TfuDkk9) {% endstep %} {% endstepper %} #### 📂 Container Structure \* \`/workspace/work/\` — Your mounted work directory \* \`/workspace/unsloth-notebooks/\` — Example fine-tuning notebooks \* \`/home/unsloth/\` — User home directory ### 📖 Usage Example #### Full Example \`\`\`bash docker run -d -e JUPYTER\_PORT=8000 \\ -e JUPYTER\_PASSWORD="mypassword" \\ -e "SSH\_KEY=$(cat ~/.ssh/container\_key.pub)" \\ -e USER\_PASSWORD="unsloth2024" \\ -p 8000:8000 -p 2222:22 \\ -v $(pwd)/work:/workspace/work \\ --gpus all \\ unsloth/unsloth \`\`\` #### Setting up SSH Key If you don't have an SSH key pair: \`\`\`bash # Generate new key pair ssh-keygen -t rsa -b 4096 -f ~/.ssh/container\_key # Use the public key in docker run -e "SSH\_KEY=$(cat ~/.ssh/container\_key.pub)" # Connect via SSH ssh -i ~/.ssh/container\_key -p 2222 unsloth@localhost \`\`\` ### 🦥Why Unsloth Containers? \* \*\*Reliable\*\*: Curated environment with stable & maintained package versions. Just 7 GB compressed (vs. 10–11 GB elsewhere) \* \*\*Ready-to-use\*\*: Pre-installed notebooks in \`/workspace/unsloth-notebooks/\` \* \*\*Secure\*\*: Runs safely as a non-root user \* \*\*Universal\*\*: Compatible with all transformer-based models (TTS, BERT, etc.) ### \*\*Unsloth not detecting or using my GPU\*\* If the model is not using your GPU specifically for Docker, try: Pulling the latest image manually: \`\`\`bash docker pull unsloth/unsloth:latest \`\`\` \* Start the container with GPU access: \* \`docker run\`: \`--gpus all\` \* Docker Compose: \`capabilities: \[gpu\]\` \* On Linux, make sure the NVIDIA Container Toolkit is installed. \* On Windows: \* Check that \`nvcc --version\` matches the CUDA version shown in \`nvidia-smi\` \* Follow: \### ⚙️ Advanced Settings \`\`\`bash # Generate SSH key pair ssh-keygen -t rsa -b 4096 -f ~/.ssh/container\_key # Connect to container ssh -i ~/.ssh/container\_key -p 2222 unsloth@localhost \`\`\` | Variable | Description | Default | | ------------------ | ---------------------------------- | --------- | | \`JUPYTER\_PASSWORD\` | Jupyter Lab password | \`unsloth\` | | \`JUPYTER\_PORT\` | Jupyter Lab port inside container | \`8888\` | | \`SSH\_KEY\` | SSH public key for authentication | \`None\` | | \`USER\_PASSWORD\` | Password for \`unsloth\` user (sudo) | \`unsloth\` | \`\`\`bash -p : \`\`\` \* Jupyter Lab: \`-p 8000:8888\` \* SSH access: \`-p 2222:22\` {% hint style="warning" %} \*\*Important\*\*: Use volume mounts to preserve your work between container runs. {% endhint %} \`\`\`bash -v : \`\`\` \`\`\`bash docker run -d -e JUPYTER\_PORT=8000 \\ -e JUPYTER\_PASSWORD="mypassword" \\ -e "SSH\_KEY=$(cat ~/.ssh/container\_key.pub)" \\ -e USER\_PASSWORD="unsloth2024" \\ -p 8000:8000 -p 2222:22 \\ -v $(pwd)/work:/workspace/work \\ --gpus all \\ unsloth/unsloth \`\`\` ### \*\*🔒 Security Notes\*\* \* Container runs as non-root \`unsloth\` user by default \* Use \`USER\_PASSWORD\` for sudo operations inside container \* SSH access requires public key authentication --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/docker.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/install/amd.md). # Guide de fine-tuning des LLM sur GPU AMD avec Unsloth Affinez les LLM jusqu’à 2× plus vite avec \\~70 % de mémoire en moins sur matériel AMD, sans NVIDIA requis. Unsloth prend en charge AMD Radeon RDNA 3/3.5/4 (séries RX 6000–9000) sur Windows et Linux, ainsi que les GPU pour centres de données, y compris le MI300X (192 Go). {% stepper %} {% step %} #### \*\*Installateur en une ligne\*\* \*\*Installation la plus simple :\*\* Passez toutes les étapes ci-dessous avec l’installateur en une ligne : il détecte automatiquement votre GPU AMD, installe PyTorch optimisé pour ROCm, bitsandbytes et lance Unsloth Studio : \*\*Linux :\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows (PowerShell) :\*\* \`\`\`powershell irm https://unsloth.ai/install.ps1 | iex \`\`\` Les étapes manuelles ci-dessous s’adressent aux utilisateurs qui souhaitent installer la bibliothèque Python Unsloth sur AMD avec les dépendances requises. {% endstep %} {% step %} #### \*\*Créer un nouvel environnement isolé (facultatif)\*\* Pour ne pas casser les paquets système, vous pouvez créer un environnement pip isolé. N’oubliez pas de vérifier la version de Python que vous avez ! Cela peut être \`pip3\`, \`pip3.13\`, \`python3\`, \`python.3.13\` etc. \*\*Linux :\*\* 3.13 indiqué ; toute version 3.11-3.13 fonctionne partout (3.10 fonctionne pour les installations manuelles {% code overflow="wrap" %} \`\`\`bash # Linux — remplacez 3.13 par la version 3.10-3.13 que vous avez apt update && apt install python3.13-venv -y python3.13 -m venv unsloth\_env source unsloth\_env/bin/activate pip install uv \`\`\` {% endcode %} \*\*Windows (PowerShell) :\*\* 3.13 indiqué ; utilisez 3.12 si vous installez l’option supplémentaire unsloth\\\[rocm72-torch291\] \`\`\`shellscript py -3.13 -m venv unsloth\_env unsloth\_env\\Scripts\\Activate.ps1 pip install uv \`\`\` {% endstep %} {% step %} #### \*\*Installer PyTorch\*\* Vous prévoyez d’installer Unsloth avec une option AMD (\`unsloth\[rocm72-torch291\]\` etc., voir la section suivante) ? Elles incluent un PyTorch correspondant, vous pouvez donc \*\*ignorer cette section\*\*. Installez PyTorch ici uniquement si vous souhaitez le gérer vous-même ou si votre version de ROCm n’est pas couverte par une option. \*\*Linux :\*\* \\ Installez PyTorch avec la prise en charge de ROCm depuis l’index PyTorch. Vérifiez votre version de ROCm via \`amd-smi version\` (repérez la ligne \`ROCm version:\` ), puis remplacez \`https://download.pytorch.org/whl/rocm7.1\` pour qu’elle corresponde. ROCm 6.0 ou plus récent est requis. {% code overflow="wrap" %} \`\`\`bash uv pip install "torch>=2.4,<2.11.0" "torchvision<0.26.0" "torchaudio<2.11.0" \\ --index-url https://download.pytorch.org/whl/rocm7.1 --upgrade --force-reinstall \`\`\` {% endcode %} ROCm 7.2 fournit des wheels plus récentes (torch 2.11), donc sous ROCm 7.2 utilisez plutôt ceci : \`\`\`bash uv pip install "torch>=2.11.0,<2.12.0" torchvision torchaudio \\ --index-url https://download.pytorch.org/whl/rocm7.2 --upgrade --force-reinstall \`\`\` Les tags d’index disponibles sont \`rocm6.0\`, \`rocm6.1\`, \`rocm6.2\`, \`rocm6.3\`, \`rocm6.4\`, \`rocm7.0\`, \`rocm7.1\`, et \`rocm7.2\`. ROCm 6.5-6.9 n’a pas de wheels dédiées (utilisez \`rocm6.4\`), et ROCm 7.3+ utilise \`rocm7.2\`. \*Les limites de version empêchent de récupérer par erreur torch 2.11+, qui n’a des wheels que pour ROCm 7.2 et cassera les choses. Mettez à jour \`rocm7.1\` pour correspondre à votre version détectée, comme précédemment.\* Nous avons aussi écrit une seule commande de terminal pour extraire la bonne version de ROCM si cela peut vous aider. \`\`\`bash ROCM\_TAG="$({ command -v amd-smi >/dev/null 2>&1 && amd-smi version 2>/dev/null | awk -F'ROCm version: ' 'NF>1{split($2,a,"."); print "rocm"a\[1\]"."a\[2\]; ok=1; exit} END{exit !ok}'; } || { \[ -r /opt/rocm/.info/version \] && awk -F. '{print "rocm"$1"."$2; exit}' /opt/rocm/.info/version; } || { command -v hipconfig >/dev/null 2>&1 && hipconfig --version 2>/dev/null | awk -F': \*' '/HIP version/{split($2,a,"."); print "rocm"a\[1\]"."a\[2\]; ok=1; exit} END{exit !ok}'; } || { command -v dpkg-query >/dev/null 2>&1 && ver="$(dpkg-query -W -f="${Version}\\n" rocm-core 2>/dev/null)" && \[ -n "$ver" \] && awk -F'\[.-\]' '{print "rocm"$1"."$2; exit}' <<<"$ver"; } || { command -v rpm >/dev/null 2>&1 && ver="$(rpm -q --qf '%{VERSION}\\n' rocm-core 2>/dev/null)" && \[ -n "$ver" \] && awk -F'\[.-\]' '{print "rocm"$1"."$2; exit}' <<<"$ver"; })"; \[ -n "$ROCM\_TAG" \] && uv pip install "torch>=2.4,<2.11.0" "torchvision<0.26.0" "torchaudio<2.11.0" --index-url "https://download.pytorch.org/whl/$ROCM\_TAG" --upgrade --force-reinstall \`\`\` \*Remarque : si votre version de ROCm est 7.2 ou supérieure, remplacez \`$ROCM\_TAG\` dans la commande ci-dessus par \`rocm7.1,\` aucune wheel PyTorch n’existe encore pour 7.2+.\* ![](https://unsloth.ai/files/e882b3336a5abe59d65f49ffa24fde2c79c9c8a7) {% endstep %} {% step %} #### \*\*Installer Unsloth\*\* Installez Unsloth avec les options AMD : {% code overflow="wrap" %} \`\`\`bash uv pip install unsloth\[amd\] \`\`\` {% endcode %} ![](https://unsloth.ai/files/56290c9ebec3af5e28230609ed40c5e3b789b4ca) ⚠️ Requis pour AMD : installez bitsandbytes compatible ROCm\\ Tous les systèmes ROCm ont besoin d’une build préliminaire de bitsandbytes ; les versions ≤ 0.49.2 ont un bug NaN de décodage en 4 bits sur tous les GPU AMD. Remarque : utilisez \`pip\` pas \`uv\` pour cette étape, \`uv\` rejette la wheel préliminaire en raison d’une incompatibilité de version dans le nom du fichier. {% code overflow="wrap" %} \`\`\`bash # Systèmes x86\_64 : pip install --force-reinstall --no-cache-dir --no-deps \\ "https://github.com/bitsandbytes-foundation/bitsandbytes/releases/download/continuous-release\_main/bitsandbytes-1.33.7.preview-py3-none-manylinux\_2\_24\_x86\_64.whl" # Systèmes aarch64 : remplacez x86\_64 par aarch64 dans l’URL ci-dessus # Solution de repli si l’URL est inaccessible : # pip install --force-reinstall --no-cache-dir --no-deps "bitsandbytes>=0.49.1" \`\`\` {% endcode %} ![](https://unsloth.ai/files/dba5bd6e91537714f4f1978da3224b6f85d732cd) {% endstep %} {% step %} #### \*\*Commencez l’affinage avec Unsloth !\*\* Et voilà. Essayez quelques exemples dans notre page \[\*\*Unsloth Notebooks\*\*\](/docs/fr/commencer/unsloth-notebooks.md) ! Vous pouvez consulter nos guides dédiés \[d’affinage\](/docs/fr/commencer/fine-tuning-llms-guide.md) ou \[d’apprentissage par renforcement\](/docs/fr/commencer/reinforcement-learning-rl-guide.md) . Voici aussi un bref exemple : \*\*1. Définir les variables d’environnement\*\* {% code overflow="wrap" %} \`\`\`bash export HSA\_OVERRIDE\_GFX\_VERSION=9.4.2 # Requis pour AMD MI300X export HF\_HUB\_DISABLE\_XET=1 # Corrige les problèmes de téléchargement HuggingFace sur AMD \`\`\` {% endcode %} \*\*\*Remarque :\*\*\* \*\`HSA\_OVERRIDE\_GFX\_VERSION=9.4.2\` indique à ROCm de traiter votre GPU comme gfx942 (MI300X). Sans cela, certains kernels peuvent échouer à la compilation ou à l’exécution.\* \*\*2. Charger et configurer le modèle\*\* {% code overflow="wrap" %} \`\`\`python from unsloth import FastModel model, tokenizer = FastModel.from\_pretrained( model\_name = "unsloth/gemma-4-26b-a4b-it", max\_seq\_length = 2048, load\_in\_4bit = True, ) model = FastModel.get\_peft\_model( model, r = 16, lora\_alpha = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj"\], ) \`\`\` {% endcode %} \*\*3. Entraîner\*\* {% code overflow="wrap" %} \`\`\`python from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( model = model, tokenizer = tokenizer, train\_dataset = dataset, formatting\_func = formatting\_func, args = SFTConfig( per\_device\_train\_batch\_size = 1, gradient\_accumulation\_steps = 4, max\_steps = 60, output\_dir = "outputs", report\_to = "none", ), ) trainer\_stats = trainer.train() \`\`\` {% endcode %} ![](https://unsloth.ai/files/7fc28a6dd0df02da2d8263b7a5153dc763f9f25b) \*\*\*Remarque :\*\* Sur les GPU AMD, Flash Attention 2 n’est pas disponible. Unsloth revient automatiquement à Xformers, qui offre des performances équivalentes sur ROCm. L’avertissement peut être ignoré en toute sécurité.\* {% endstep %} {% endstepper %} ### :1234: Apprentissage par renforcement sur GPU AMD Vous pouvez utiliser notre exemple :ledger:\[gpt-oss RL auto win 2048\](https://github.com/unslothai/notebooks/blob/main/nb/AMD-gpt\_oss\_\\(20B\\)\_Reinforcement\_Learning\_2048\_Game\_BF16.ipynb) sur un GPU MI300X (192 Go). L’objectif est de jouer automatiquement au jeu 2048 et de le gagner avec RL. Le LLM (gpt-oss 20b) élabore automatiquement une stratégie pour gagner au jeu 2048, et nous attribuons une récompense élevée aux stratégies gagnantes, et de faibles récompenses aux stratégies perdantes. {% columns %} {% column %} ![](https://unsloth.ai/files/3cadd91340e63b0562100bb73710efc5c7359901) {% endcolumn %} {% column %} La récompense au fil du temps augmente après environ 300 étapes environ ! L’objectif du RL est de maximiser la récompense moyenne pour gagner au jeu 2048. ![](https://unsloth.ai/files/35798f90e1ed33ff069b0c1ce1a8444b567044a1) {% endcolumn %} {% endcolumns %} Nous avons utilisé une machine AMD MI300X (192 Go) pour exécuter l’exemple RL 2048 avec Unsloth, et cela a très bien fonctionné ! ![](https://unsloth.ai/files/204d29be924da780008880c800f066cd8e3f4ee4) ![](https://unsloth.ai/files/8a689b36bb45d149bc645e55f346f590928afe54) Vous pouvez également utiliser notre :ledger:\[notebook RL de génération automatique de kernels\](https://github.com/unslothai/notebooks/blob/main/nb/AMD-gpt\_oss\_\\(20B\\)\_GRPO\_BF16.ipynb) également avec gpt-oss pour créer automatiquement des kernels de multiplication matricielle en Python. Le notebook conçoit aussi plusieurs méthodes pour contrer le reward hacking. {% columns %} {% column width="50%" %} L’invite que nous avons utilisée pour créer automatiquement ces kernels était : {% code overflow="wrap" %} \`\`\`\` Créez une nouvelle fonction de multiplication matricielle rapide en utilisant uniquement du code Python natif. Vous recevez une liste de listes de nombres. Affichez votre nouvelle fonction entre des backticks en utilisant le format ci-dessous : \`\`\` python def matmul(A, B): return ... \`\`\` \`\`\`\` {% endcode %} {% endcolumn %} {% column width="50%" %} Le processus de RL apprend par exemple à appliquer l’algorithme de Strassen pour accélérer la multiplication matricielle en Python. ![](https://unsloth.ai/files/c3c38d8a05f5c9a2f3e29ff5012d3bf334fb1a8e) {% endcolumn %} {% endcolumns %} ### :books:Notebooks gratuits AMD en un clic AMD fournit des notebooks en un clic équipés de \*\*GPU MI300X gratuits avec 192 Go de VRAM\*\* via leur Dev Cloud. Entraînez de grands modèles entièrement gratuitement (aucune inscription ni carte de crédit requise) : \* \[Qwen3 (32B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(32B\\)\_A100-Reasoning-Conversational.ipynb) \* \[Llama 3.3 (70B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Llama3.3\_\\(70B\\)\_A100-Conversational.ipynb) \* \[Qwen3 (14B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* \[Mistral v0.3 (7B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Mistral\_v0.3\_\\(7B\\)-Alpaca.ipynb) \* \[GPT OSS MXFP4 (20B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-GPT\_OSS\_MXFP4\_\\(20B\\)-Inference.ipynb) - Inférence \* \[Gemma4 (E2B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) - RL Sudoku \* Unsloth Studio {% embed url="" %} Vous pouvez utiliser n’importe quel notebook Unsloth en le faisant précéder de dans \[Notebooks Unsloth\](/docs/fr/commencer/unsloth-notebooks.md) en modifiant le lien de \\ à {% columns %} {% column width="33.33333333333333%" %} ![](https://unsloth.ai/files/8538e846c2aeb1caa27f6be7a045669eb329ef82) {% endcolumn %} {% column width="66.66666666666667%" %} ![](https://unsloth.ai/files/42e3a49219a1f0ee0fab09cbf2cc45838c810c4c) {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/install/amd.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/pip-install.md). # Install Unsloth via pip and uv Unsloth can be used in two ways: through \[Unsloth Studio\](#unsloth-studio), the web UI, or through \[Unsloth Core\](#unsloth-core), the code-based version. ## \*\*Unsloth Studio\*\* #### \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Use the same command to update. #### \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` Use the same command to update. #### Launch: \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` For detailed Unsloth Studio install instructions and requirements, \[view our guide\](/docs/new/studio/install.md). ### \*\*Install from Main Repo\*\* #### \*\*macOS, Linux, WSL developer installs:\*\* \`\`\`bash git clone https://github.com/unslothai/unsloth cd unsloth ./install.sh --local unsloth studio -H 0.0.0.0 -p 8888 \`\`\` #### \*\*Windows PowerShell developer installs:\*\* \`\`\`powershell winget install -e --id Python.Python.3.13 --source winget winget install --id=astral-sh.uv -e --source winget winget install --id Git.Git -e --source winget git clone https://github.com/unslothai/unsloth cd unsloth Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass .\\install.ps1 --local unsloth studio -H 0.0.0.0 -p 8888 \`\`\` ### \*\*Nightly Install\*\* #### \*\*Nightly - MacOS, Linux, WSL:\*\* \`\`\`bash git clone https://github.com/unslothai/unsloth cd unsloth git checkout nightly ./install.sh --local \`\`\` Then to launch every time: \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` #### \*\*Nightly - Windows:\*\* Run in Windows Powershell: \`\`\`bash winget install -e --id Python.Python.3.13 --source winget winget install --id=astral-sh.uv -e --source winget winget install --id Git.Git -e --source winget git clone https://github.com/unslothai/unsloth cd unsloth git checkout nightly Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass .\\install.ps1 --local \`\`\` Then to launch every time: unsloth studio -H 0.0.0.0 -p 8888 \### Uninstall The recommended way to fully remove Unsloth Studio is the uninstall script for your OS. It stops any running servers, removes the app, CLI command, launcher data, shortcuts, and platform-specific entries (macOS \`.app\` bundle + Launch Services; Windows Start Menu + registry + PATH): \*\*macOS, WSL, Linux:\*\* \`\`\`shellscript curl -fsSL https://raw.githubusercontent.com/unslothai/unsloth/main/scripts/uninstall.sh | sh \`\`\` \*\*Windows (PowerShell):\*\* \`\`\`ps1 irm https://raw.githubusercontent.com/unslothai/unsloth/main/scripts/uninstall.ps1 | iex \`\`\` #### Manual uninstall If you prefer to remove only specific parts: \*\*1. Remove app only\*\* (keeps history, chats, checkpoints, and exports intact): \* \*\*macOS, WSL, Linux:\*\* \`rm -rf ~/.unsloth/studio/unsloth\_studio\` \* \*\*Windows (PowerShell):\*\* \`Remove-Item -Recurse -Force "$HOME\\.unsloth\\studio\\unsloth\_studio"\` \*\*2. Remove Unsloth entirely\*\* (keeps other Unsloth tools intact): \* \*\*macOS, WSL, Linux:\*\* \`rm -rf ~/.unsloth/studio\` \* \*\*Windows (PowerShell):\*\* \`Remove-Item -Recurse -Force "$HOME\\.unsloth\\studio"\` \*\*3. Remove everything Unsloth-related:\*\* \* \*\*macOS, WSL, Linux:\*\* \`rm -rf ~/.unsloth\` \* \*\*Windows (PowerShell):\*\* \`Remove-Item -Recurse -Force "$HOME\\.unsloth"\` {% hint style="warning" %} Note: Step 3 deletes everything in history, chats, model checkpoints, and exports. This cannot be undone. {% endhint %} \*\*4. Remove shortcuts and symlinks:\*\* \*\*macOS:\*\* \`\`\`shellscript rm -rf ~/Applications/Unsloth\\ Studio.app ~/Desktop/Unsloth\\ Studio \`\`\` \*\*Linux:\*\* \`\`\`shellscript rm -f ~/.local/share/applications/unsloth-studio.desktop ~/Desktop/unsloth-studio.desktop \`\`\` \*\*WSL / Windows (PowerShell):\*\* \`\`\`shellscript Remove-Item -Force "$HOME\\Desktop\\Unsloth Studio.lnk" Remove-Item -Force "$env:APPDATA\\Microsoft\\Windows\\Start Menu\\Programs\\Unsloth Studio.lnk" \`\`\` \*\*5. Remove the CLI command:\*\* \* \*\*macOS, Linux, WSL:\*\* \`rm -f ~/.local/bin/unsloth\` \* \*\*Windows (PowerShell):\*\* The installer added the venv's \`Scripts\` directory to your User PATH. To remove it, open Settings → System → About → Advanced system settings → Environment Variables, find \`Path\` under User variables, and remove the entry pointing to \`.unsloth\\studio\\...\\Scripts\`. {% hint style="info" %} Note: Steps 1-5 dont touch your downloaded HF model files. See Deleting cached HF model files below if you want to reclaim that space. {% endhint %} ### \*\*Deleting cached HF model files\*\* You can delete old model files either from the bin icon in model search or by removing the relevant cached model folder from the default Hugging Face cache directory. By default, Hugging Face uses \`~/.cache/huggingface/hub/\` on macOS/Linux/WSL and \`C:\\Users\\\\.cache\\huggingface\\hub\\\` on Windows. \* \*\*MacOS, Linux, WSL:\*\* \`~/.cache/huggingface/hub/\` \* \*\*Windows:\*\* \`%USERPROFILE%\\.cache\\huggingface\\hub\\\` If \`HF\_HUB\_CACHE\` or \`HF\_HOME\` is set, use that location instead. On Linux and WSL, \`XDG\_CACHE\_HOME\` can also change the default cache root. ## \*\*Unsloth Core\*\* \*\*Install uv with the following command:\*\* \*\*macOS, WSL, Linux:\*\* \`\`\`shellscript curl -LsSf https://astral.sh/uv/install.sh | sh \`\`\` \*\*Windows (PowerShell):\*\* \`\`\`powershell irm https://astral.sh/uv/install.ps1 | iex \`\`\` \*\*Install unsloth core with uv pip (recommended) for the latest pip release:\*\* \`\`\`bash uv venv unsloth\_env --python 3.13 source unsloth\_env/bin/activate uv pip install unsloth --torch-backend=auto \`\`\` Or just pip: \`\`\`bash pip install unsloth \`\`\` To install \*\*vLLM and Unsloth\*\* together, do: \`\`\`bash uv pip install unsloth vllm --torch-backend=auto \`\`\` To install the \*\*latest main branch\*\* of Unsloth, do: {% code overflow="wrap" %} \`\`\`bash uv pip install unsloth --torch-backend=auto pip uninstall unsloth unsloth\_zoo -y && pip install --no-deps git+https://github.com/unslothai/unsloth\_zoo.git && pip install --no-deps git+https://github.com/unslothai/unsloth.git \`\`\` {% endcode %} For \*\*venv and virtual environments installs\*\* to isolate your installation to not break system packages, and to reduce irreparable damage to your system, use venv: {% code overflow="wrap" %} \`\`\`bash apt install python3.10-venv python3.11-venv python3.12-venv python3.13-venv -y python -m venv unsloth\_env source unsloth\_env/bin/activate pip install --upgrade pip && pip install uv uv pip install unsloth --torch-backend=auto \`\`\` {% endcode %} If you're installing Unsloth in Jupyter, Colab, or other notebooks, be sure to prefix the command with \`!\`. This isn't necessary when using a terminal {% hint style="info" %} Python 3.13 is now supported! {% endhint %} ### Uninstall Unsloth Core If you're still encountering dependency issues with Unsloth, many users have resolved them by forcing uninstalling and reinstalling Unsloth: {% code overflow="wrap" %} \`\`\`bash pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth\_zoo \`\`\` {% endcode %} ### Advanced Pip Installation {% hint style="warning" %} Do \*\*NOT\*\* use this if you have \[Conda\](/docs/get-started/install/conda-install.md). {% endhint %} Pip is a bit more complex since there are dependency issues. The pip command is different for \`torch 2.2,2.3,2.4,2.5\` and CUDA versions. For other torch versions, we support \`torch211\`, \`torch212\`, \`torch220\`, \`torch230\`, \`torch240\` and for CUDA versions, we support \`cu118\` and \`cu121\` and \`cu124\`. For Ampere devices (A100, H100, RTX3090) and above, use \`cu118-ampere\` or \`cu121-ampere\` or \`cu124-ampere\`. For example, if you have \`torch 2.4\` and \`CUDA 12.1\`, use: \`\`\`bash pip install --upgrade pip pip install "unsloth\[cu121-torch240\] @ git+https://github.com/unslothai/unsloth.git" \`\`\` Another example, if you have \`torch 2.5\` and \`CUDA 12.4\`, use: \`\`\`bash pip install --upgrade pip pip install "unsloth\[cu124-torch250\] @ git+https://github.com/unslothai/unsloth.git" \`\`\` And other examples: \`\`\`bash pip install "unsloth\[cu121-ampere-torch240\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu118-ampere-torch240\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu121-torch240\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu118-torch240\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu121-torch230\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu121-ampere-torch230\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu121-torch250\] @ git+https://github.com/unslothai/unsloth.git" pip install "unsloth\[cu124-ampere-torch250\] @ git+https://github.com/unslothai/unsloth.git" \`\`\` Or, run the below in a terminal to get the \*\*optimal\*\* pip installation command: \`\`\`bash wget -qO- https://raw.githubusercontent.com/unslothai/unsloth/main/unsloth/\_auto\_install.py | python - \`\`\` Or, run the below manually in a Python REPL: {% code overflow="wrap" %} \`\`\`python # Licensed under the Apache License, Version 2.0 (the "License") try: import torch except: raise ImportError('Install torch via \`pip install torch\`') from packaging.version import Version as V import re v = V(re.match(r"\[0-9\\.\]{3,}", torch.\_\_version\_\_).group(0)) cuda = str(torch.version.cuda) is\_ampere = torch.cuda.get\_device\_capability()\[0\] >= 8 USE\_ABI = torch.\_C.\_GLIBCXX\_USE\_CXX11\_ABI if cuda not in ("11.8", "12.1", "12.4", "12.6", "12.8", "13.0"): raise RuntimeError(f"CUDA = {cuda} not supported!") if v <= V('2.1.0'): raise RuntimeError(f"Torch = {v} too old!") elif v <= V('2.1.1'): x = 'cu{}{}-torch211' elif v <= V('2.1.2'): x = 'cu{}{}-torch212' elif v < V('2.3.0'): x = 'cu{}{}-torch220' elif v < V('2.4.0'): x = 'cu{}{}-torch230' elif v < V('2.5.0'): x = 'cu{}{}-torch240' elif v < V('2.5.1'): x = 'cu{}{}-torch250' elif v <= V('2.5.1'): x = 'cu{}{}-torch251' elif v < V('2.7.0'): x = 'cu{}{}-torch260' elif v < V('2.7.9'): x = 'cu{}{}-torch270' elif v < V('2.8.0'): x = 'cu{}{}-torch271' elif v < V('2.8.9'): x = 'cu{}{}-torch280' elif v < V('2.9.1'): x = 'cu{}{}-torch290' elif v < V('2.9.2'): x = 'cu{}{}-torch291' else: raise RuntimeError(f"Torch = {v} too new!") if v > V('2.6.9') and cuda not in ("11.8", "12.6", "12.8", "13.0"): raise RuntimeError(f"CUDA = {cuda} not supported!") x = x.format(cuda.replace(".", ""), "-ampere" if False else "") # is\_ampere is broken due to flash-attn print(f'pip install --upgrade pip && pip install --no-deps git+https://github.com/unslothai/unsloth-zoo.git && pip install "unsloth\[{x}\] @ git+https://github.com/unslothai/unsloth.git" --no-build-isolation') \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/pip-install.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/sglang-guide.md). # SGLang Deployment & Inference Guide You can serve any LLM or fine-tuned model via \[SGLang\](https://github.com/sgl-project/sglang) for low-latency, high-throughput inference. SGLang supports text, image/video model inference on any GPU setup, with support for some GGUFs. ### :computer:Installing SGLang To install SGLang and Unsloth on NVIDIA GPUs, you can use the below in a virtual environment (which won't break your other Python libraries) \`\`\`shellscript # OPTIONAL use a virtual environment python -m venv unsloth\_env source unsloth\_env/bin/activate # Install Rust, outlines-core then SGLang curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh source $HOME/.cargo/env && sudo apt-get install -y pkg-config libssl-dev pip install --upgrade pip && pip install uv uv pip install "sglang" && uv pip install unsloth \`\`\` For \*\*Docker\*\* setups run: {% code overflow="wrap" %} \`\`\`shellscript docker run --gpus all \\ --shm-size 32g \\ -p 30000:30000 \\ -v ~/.cache/huggingface:/root/.cache/huggingface \\ --env "HF\_TOKEN=" \\ --ipc=host \\ lmsysorg/sglang:latest \\ python3 -m sglang.launch\_server --model-path unsloth/Llama-3.1-8B-Instruct --host 0.0.0.0 --port 30000 \`\`\` {% endcode %} ### :bug:Debugging SGLang Installation issues Note if you see the below, update Rust and outlines-core as specified in \[#setting-up-sglang\](#setting-up-sglang "mention") {% code overflow="wrap" %} \`\`\` hint: This usually indicates a problem with the package or the build environment. help: \`outlines-core\` (v0.1.26) was included because \`sglang\` (v0.5.5.post2) depends on \`outlines\` (v0.1.11) which depends on \`outlines-core\` \`\`\` {% endcode %} If you see a Flashinfer issue like below: \`\`\` /home/daniel/.cache/flashinfer/0.5.2/100a/generated/batch\_prefill\_with\_kv\_cache\_dtype\_q\_bf16\_dtype\_kv\_bf16\_dtype\_o\_bf16\_dtype\_idx\_i32\_head\_dim\_qk\_64\_head\_dim\_vo\_64\_posenc\_0\_use\_swa\_False\_use\_logits\_cap\_False\_f16qk\_False/batch\_prefill\_ragged\_kernel\_mask\_1.cu:1:10: fatal error: flashinfer/attention/prefill.cuh: No such file or directory 1 | #include | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ compilation terminated. ninja: build stopped: subcommand failed. Possible solutions: 1. set --mem-fraction-static to a smaller value (e.g., 0.8 or 0.7) 2. set --cuda-graph-max-bs to a smaller value (e.g., 16) 3. disable torch compile by not using --enable-torch-compile 4. disable CUDA graph by --disable-cuda-graph. (Not recommended. Huge performance loss) Open an issue on GitHub https://github.com/sgl-project/sglang/issues/new/choose \`\`\` Remove the flashinfer cache via \`rm -rf .cache/flashinfer\` and also the directory listed in the error message ie \`rm -rf ~/.cache/flashinfer\` ### :truck:Deploying SGLang models To deploy any model like for example \[unsloth/Llama-3.2-1B-Instruct\](https://huggingface.co/unsloth/Llama-3.2-1B-Instruct), do the below in a separate terminal (otherwise it'll block your current terminal - you can also use tmux): {% code overflow="wrap" %} \`\`\`shellscript python3 -m sglang.launch\_server \\ --model-path unsloth/Llama-3.2-1B-Instruct \\ --host 0.0.0.0 --port 30000 \`\`\` {% endcode %} ![](https://unsloth.ai/files/RGebdoyQToYxdn7ndUMx) You can then use the OpenAI Chat completions library to call the model (in another terminal or using tmux): \`\`\`python # Install openai via pip install openai from openai import OpenAI import json openai\_client = OpenAI( base\_url = "http://0.0.0.0:30000/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/Llama-3.2-1B-Instruct", messages = \[{"role": "user", "content": "What is 2+2?"},\], ) print(completion.choices\[0\].message.content) \`\`\` And you will get \`2 + 2 = 4.\` ### 🦥Deploying Unsloth finetunes in SGLang After fine-tuning \[Fine-tuning Guide\](/docs/get-started/fine-tuning-llms-guide.md) or using our notebooks at \[Unsloth Notebooks\](/docs/get-started/unsloth-notebooks.md), you can save or deploy your models directly through SGLang within a single workflow. An example Unsloth finetuning script for eg: \`\`\`python from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/gpt-oss-20b", max\_seq\_length = 2048, load\_in\_4bit = True, ) model = FastLanguageModel.get\_peft\_model(model) \`\`\` \*\*To save to 16-bit for SGLang, use:\*\* \`\`\`python model.save\_pretrained\_merged("finetuned\_model", tokenizer, save\_method = "merged\_16bit") ## OR to upload to HuggingFace: model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "merged\_16bit", token = "") \`\`\` \*\*To save just the LoRA adapters\*\*, either use: \`\`\`python model.save\_pretrained("finetuned\_model") tokenizer.save\_pretrained("finetuned\_model") \`\`\` Or just use our builtin function to do that: \`\`\`python model.save\_pretrained\_merged("model", tokenizer, save\_method = "lora") ## OR to upload to HuggingFace model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "lora", token = "") \`\`\` ### :railway\\\_car:gpt-oss-20b: Unsloth & SGLang Deployment Guide Below is a step-by-step tutorial with instructions for training the \[gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune.md)-20b using Unsloth and deploying it with SGLang. It includes performance benchmarks across multiple quantization formats. {% stepper %} {% step %} #### Unsloth Fine-tuning and Exporting Formats If you're new to fine-tuning, you can read our \[guide\](/docs/get-started/fine-tuning-llms-guide.md), or try the gpt-oss 20B finetuning notebook at \[gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune.md) After training, you can export the model in multiple formats: {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged( "finetuned\_model", tokenizer, save\_method = "merged\_16bit", ) ## For gpt-oss specific mxfp4 conversions: model.save\_pretrained\_merged( "finetuned\_model", tokenizer, save\_method = "mxfp4", # (ONLY FOR gpt-oss otherwise choose "merged\_16bit") ) \`\`\` {% endcode %} {% endstep %} {% step %} #### Deployment with SGLang We saved our gpt-oss finetune to the folder "finetuned\\\_model", and so in a new terminal, we can launch the finetuned model as an inference endpoint with SGLang: \`\`\`shellscript python -m sglang.launch\_server \\ --model-path finetuned\_model \\ --host 0.0.0.0 --port 30002 \`\`\` You might have to wait a bit on \`Capturing batches (bs=1 avail\_mem=20.84 GB):\` ! {% endstep %} {% step %} #### Calling the inference endpoint: To call the inference endpoint, first launch a new terminal. We then can call the model like below: {% code overflow="wrap" %} \`\`\`python from openai import OpenAI import json openai\_client = OpenAI( base\_url = "http://0.0.0.0:30002/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "finetuned\_model", messages = \[{"role": "user", "content": "What is 2+2?"},\], ) print(completion.choices\[0\].message.content) ## OUTPUT ## # <|channel|>analysis<|message|>The user asks a simple math question. We should answer 4. Also we should comply with policy. No issues.<|end|><|start|>assistant<|channel|>final<|message|>2 + 2 equals 4. \`\`\` {% endcode %} {% endstep %} {% endstepper %} ### :gem:FP8 Online Quantization To deploy models with FP8 online quantization which allows 30 to 50% more throughput and 50% less memory usage with 2x longer context length supports with SGLang, you can do the below: {% code overflow="wrap" %} \`\`\`shellscript python -m sglang.launch\_server \\ --model-path unsloth/Llama-3.2-1B-Instruct \\ --host 0.0.0.0 --port 30002 \\ --quantization fp8 \\ --kv-cache-dtype fp8\_e4m3 \`\`\` {% endcode %} You can also use \`--kv-cache-dtype fp8\_e5m2\` which has a larger dynamic range which might solve FP8 inference issues if you see them. Or use our pre-quantized float8 quants listed in or some are listed below: {% embed url="" %} {% embed url="" %} ### ⚡Benchmarking SGLang Below is some code you can run to test the performance speed of your finetuned model: \`\`\`shellscript python -m sglang.launch\_server \\ --model-path finetuned\_model \\ --host 0.0.0.0 --port 30002 \`\`\` Then in another terminal or via tmux: \`\`\`shellscript # Batch Size=8, Input=1024, Output=1024 python -m sglang.bench\_one\_batch\_server \\ --model finetuned\_model \\ --base-url http://0.0.0.0:30002 \\ --batch-size 8 \\ --input-len 1024 \\ --output-len 1024 \`\`\` You will see the benchmarking run like below: ![](https://unsloth.ai/files/BXlmCZ2892ssh6Ra9oKu) We used a B200x1 GPU with gpt-oss-20b and got the below results (\\~2,500 tokens throughput) | Batch/Input/Output | TTFT (s) | ITL (s) | Input Throughput | Output Throughput | | ------------------ | -------- | ------- | ---------------- | ----------------- | | 8/1024/1024 | 0.40 | 3.59 | 20,718.95 | 2,562.87 | | 8/8192/1024 | 0.42 | 3.74 | 154,459.01 | 2,473.84 | See for server arguments for SGLang. ### :person\\\_running:SGLang Interactive Offline Mode You can also use SGLang in offline mode (ie not a server) inside a Python interactive environment. {% code overflow="wrap" %} \`\`\`python import sglang as sgl engine = sgl.Engine(model\_path = "unsloth/Qwen3-0.6B", random\_seed = 42) prompt = "Today is a sunny day and I like" sampling\_params = {"temperature": 0, "max\_new\_tokens": 256} outputs = engine.generate(prompt, sampling\_params)\["text"\] print(outputs) engine.shutdown() \`\`\` {% endcode %} ### :sparkler:GGUFs in SGLang SGLang also interestingly supports GGUFs! \*\*Qwen3 MoE is still under construction, but most dense models (Llama 3, Qwen 3, Mistral etc) are supported.\*\* First install the latest gguf python package via: {% code overflow="wrap" %} \`\`\`shellscript pip install -e "git+https://github.com/ggml-org/llama.cpp.git#egg=gguf&subdirectory=gguf-py" # install a python package from a repo subdirectory \`\`\` {% endcode %} Then for example in offline mode SGLang, you can do: {% code overflow="wrap" %} \`\`\`python from huggingface\_hub import hf\_hub\_download model\_path = hf\_hub\_download( "unsloth/Qwen3-32B-GGUF", filename = "Qwen3-32B-UD-Q4\_K\_XL.gguf", ) import sglang as sgl engine = sgl.Engine(model\_path = model\_path, random\_seed = 42) prompt = "Today is a sunny day and I like" sampling\_params = {"temperature": 0, "max\_new\_tokens": 256} outputs = engine.generate(prompt, sampling\_params)\["text"\] print(outputs) engine.shutdown() \`\`\` {% endcode %} ### :clapper:High throughput GGUF serving with SGLang First download the specific GGUF file like below: {% code overflow="wrap" %} \`\`\`python from huggingface\_hub import hf\_hub\_download hf\_hub\_download("unsloth/Qwen3-32B-GGUF", filename="Qwen3-32B-UD-Q4\_K\_XL.gguf", local\_dir=".") \`\`\` {% endcode %} Then serve the specific file \`Qwen3-32B-UD-Q4\_K\_XL.gguf\` and use \`--served-model-name unsloth/Qwen3-32B\` and also we need the HuggingFace compatible tokenizer via \`--tokenizer-path\` \`\`\`shellscript python -m sglang.launch\_server \\ --model-path Qwen3-32B-UD-Q4\_K\_XL.gguf \\ --host 0.0.0.0 --port 30002 \\ --served-model-name unsloth/Qwen3-32B \\ --tokenizer-path unsloth/Qwen3-32B \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/sglang-guide.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \# Unsloth Documentation ## 🇺🇸 English - \[Unsloth Docs\](https://unsloth.ai/docs/get-started/readme.md): Unsloth is an open-source framework for running and training LLMs. - \[Unsloth Model Catalog\](https://unsloth.ai/docs/get-started/unsloth-model-catalog.md) - \[Fine-tuning for Beginners\](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners.md) - \[Unsloth Requirements\](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements.md): Here are Unsloth's requirements including system and GPU VRAM requirements. - \[FAQ + Is Fine-tuning Right For Me?\](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me.md): If you're stuck on if fine-tuning is right for you, see here! Learn about fine-tuning misconceptions, how it compared to RAG and more: - \[Unsloth Notebooks\](https://unsloth.ai/docs/get-started/unsloth-notebooks.md): Fine-tuning notebooks: Explore the Unsloth catalog. - \[Unsloth Installation\](https://unsloth.ai/docs/get-started/install.md): Learn to install Unsloth locally or online. - \[Install Unsloth via pip and uv\](https://unsloth.ai/docs/get-started/install/pip-install.md): To install Unsloth locally via Pip, follow the steps below: - \[Install Unsloth on MacOS\](https://unsloth.ai/docs/get-started/install/mac.md) - \[How to Fine-Tune LLMs on Windows with Unsloth (Step-by-Step Guide)\](https://unsloth.ai/docs/get-started/install/windows-installation.md): See how to install Unsloth on Windows to start fine-tuning LLMs locally. - \[Fine-tuning LLMs on AMD GPUs with Unsloth Guide\](https://unsloth.ai/docs/get-started/install/amd.md): Learn how to fine-tune large language models (LLMs) on AMD GPUs with Unsloth. - \[AMD AI Reinforcement Learning Hackathon with Unsloth\](https://unsloth.ai/docs/get-started/install/amd/amd-hackathon.md): ​​Learn hands-on techniques for ​Reinforcement Learning for AI models with Unsloth from Daniel Han, the creator of Unsloth. - \[Install Unsloth via Docker\](https://unsloth.ai/docs/get-started/install/docker.md): Install Unsloth using our official Docker container - \[Updating Unsloth\](https://unsloth.ai/docs/get-started/install/updating.md): To update or use an old version of Unsloth, follow the steps below: - \[Fine-tuning LLMs on Intel GPUs with Unsloth\](https://unsloth.ai/docs/get-started/install/intel.md): Learn how to train and fine-tune large language models on Intel GPUs. - \[Conda Install\](https://unsloth.ai/docs/get-started/install/conda-install.md): To install Unsloth locally on Conda, follow the steps below: - \[How to Fine-tune LLMs in VS Code with Unsloth & Colab GPUs\](https://unsloth.ai/docs/get-started/install/vs-code.md): Guide to fine-tuning models directly in Visual Studio Code via Unsloth and Google Colab. - \[Google Colab\](https://unsloth.ai/docs/get-started/install/google-colab.md): To install and run Unsloth on Google Colab, follow the steps below: - \[Fine-tuning LLMs Guide\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide.md): Learn all the basics and best practices of fine-tuning. Beginner-friendly. - \[Datasets Guide\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/datasets-guide.md): Learn how to create & prepare a dataset for fine-tuning. - \[LoRA fine-tuning Hyperparameters Guide\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide.md): Learn step-by-step the best LLM fine-tuning settings - LoRA rank & alpha, epochs, batch size + gradient accumulation, QLoRA vs. LoRA, target modules, and more. - \[What Model Should I Use for Fine-tuning?\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use.md) - \[Tutorial: How to Finetune Llama-3 and Use In Ollama\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/tutorial-how-to-finetune-llama-3-and-use-in-ollama.md): Beginner's Guide for creating a customized personal assistant (like ChatGPT) to run locally on Ollama - \[Reinforcement Learning (RL) Guide\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide.md): Learn all about Reinforcement Learning (RL) and how to train your own DeepSeek-R1 reasoning model with Unsloth using GRPO. A complete guide from beginner to advanced. - \[Reinforcement Learning GRPO with 7x Longer Context\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context.md): Learn how Unsloth enables ultra long context RL fine-tuning. - \[Vision Reinforcement Learning (VLM RL)\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md): Train Vision/multimodal models via GRPO and RL with Unsloth! - \[FP8 Reinforcement Learning\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning.md): Train reinforcement learning (RL) and GRPO in FP8 precision with Unsloth. - \[Tutorial: Train your own Reasoning model with GRPO\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/tutorial-train-your-own-reasoning-model-with-grpo.md): Beginner's Guide to transforming a model like Llama 3.1 (8B) into a reasoning model by using Unsloth and GRPO. - \[Advanced Reinforcement Learning Documentation\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation.md): Advanced documentation settings when using Unsloth with GRPO. - \[GSPO Reinforcement Learning\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/gspo-reinforcement-learning.md): Train with GSPO (Group Sequence Policy Optimization) RL in Unsloth. - \[RL Reward Hacking\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/rl-reward-hacking.md): Learn what is Reward Hacking in Reinforcement Learning and how to counter it. - \[FP16 vs BF16 for RL\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl.md): Defeating the Training-Inference Mismatch via FP16 https://arxiv.org/pdf/2510.26788 shows how using float16 is better than bfloat16 - \[Memory Efficient RL\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl.md) - \[Preference Optimization Training - DPO, ORPO & KTO\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/preference-dpo-orpo-and-kto.md): Learn about preference alignment fine-tuning with DPO, GRPO, ORPO or KTO via Unsloth, follow the steps below: - \[Training AI Agents with RL\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/training-ai-agents-with-rl.md): Learn how to train AI agents for real-world tasks using Reinforcement Learning (RL). - \[Introducing Unsloth Studio\](https://unsloth.ai/docs/new/studio.md): Run and train AI models locally with Unsloth Studio. - \[Get started with Unsloth Studio\](https://unsloth.ai/docs/new/studio/start.md): A guide for getting started with the fine-tuning studio, data recipes, model exporting, and chat. - \[How to Run models with Unsloth Studio\](https://unsloth.ai/docs/new/studio/chat.md): Run AI models, LLMs and GGUFs locally with Unsloth Studio. - \[Unsloth Studio Installation\](https://unsloth.ai/docs/new/studio/install.md): Learn how to install Unsloth Studio on your local device. - \[Unsloth Data Recipes\](https://unsloth.ai/docs/new/studio/data-recipe.md): Learn how to create, build and edit datasets with Unsloth Studio's Data Recipes. - \[Export models with Unsloth Studio\](https://unsloth.ai/docs/new/studio/export.md): Learn how to export your safetensor or LoRA model files to GGUF or other formats. - \[Unsloth Updates\](https://unsloth.ai/docs/new/changelog.md): Unsloth Changelog for our latest releases, improvements and fixes. - \[Large language model (LLMs) Tutorials\](https://unsloth.ai/docs/models/tutorials.md) - \[Qwen3 - How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune.md): Learn to run & fine-tune Qwen3 locally with Unsloth + our Dynamic 2.0 quants - \[Qwen3-VL: How to Run Guide\](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md): Learn to fine-tune and run Qwen3-VL locally with Unsloth. - \[Qwen3-2507: Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-2507.md): Run Qwen3-30B-A3B-2507 and 235B-A22B Thinking and Instruct versions locally on your device! - \[MiniMax-M2.7 - How to Run Locally\](https://unsloth.ai/docs/models/tutorials/minimax-m27.md): Run MiniMax-M2.7 LLM locally on your own device! - \[GLM-5: How to Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/glm-5.md): Run the new GLM-5 model by Z.ai on your own local device! - \[Kimi K2.5: How to Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/kimi-k2.5.md): Guide on running Kimi-K2.5 on your own local device! - \[GLM-4.7-Flash: How To Run Locally\](https://unsloth.ai/docs/models/tutorials/glm-4.7-flash.md): Run & fine-tune GLM-4.7-Flash locally on your device! - \[Gemma 3 - How to Run Guide\](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md): How to run Gemma 3 effectively with our GGUFs on llama.cpp, Ollama, Open WebUI and how to fine-tune with Unsloth! - \[Gemma 3n: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune/gemma-3n-how-to-run-and-fine-tune.md): Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth! - \[Qwen3-Coder: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/qwen3-coder-how-to-run-locally.md): Run Qwen3-Coder-30B-A3B-Instruct and 480B-A35B locally with Unsloth Dynamic quants. - \[MiniMax-M2.5: How to Run Guide\](https://unsloth.ai/docs/models/tutorials/minimax-m25.md): Run MiniMax-M2.5 locally on your own device! - \[DeepSeek-OCR 2: How to Run & Fine-tune Guide\](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-2.md): Guide on how to run and fine-tune DeepSeek-OCR-2 locally. - \[GLM-4.7: How to Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/glm-4.7.md): A guide on how to run Z.ai GLM-4.7 model on your own local device! - \[How to Run Qwen-Image-2512 Locally in ComfyUI\](https://unsloth.ai/docs/models/tutorials/qwen-image-2512.md): Step-by-step tutorial for running Qwen-Image-2512 on your local device with ComfyUI. - \[Run Qwen-Image-2512 in stable-diffusion.cpp Tutorial\](https://unsloth.ai/docs/models/tutorials/qwen-image-2512/stable-diffusion.cpp.md): Tutorial for using Qwen-Image-2512 in stable-diffusion.cpp. - \[Devstral 2 - How to Run Guide\](https://unsloth.ai/docs/models/tutorials/devstral-2.md): Guide for local running Mistral Devstral 2 models: 123B-Instruct-2512 and Small-2-24B-Instruct-2512. - \[Ministral 3 - How to Run Guide\](https://unsloth.ai/docs/models/tutorials/ministral-3.md): Guide for Mistral Ministral 3 models, to run or fine-tune locally on your device - \[DeepSeek-OCR: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md): Guide on how to run and fine-tune DeepSeek-OCR locally. - \[Kimi K2 Thinking: Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/kimi-k2-thinking-how-to-run-locally.md): Guide on running Kimi-K2-Thinking and Kimi-K2 on your own local device! - \[GLM-4.6: Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/glm-4.6-how-to-run-locally.md): A guide on how to run Z.ai GLM-4.6 and GLM-4.6V-Flash model on your own local device! - \[Qwen3-Next: Run Locally Guide\](https://unsloth.ai/docs/models/tutorials/qwen3-next.md): Run Qwen3-Next-80B-A3B-Instruct and Thinking versions locally on your device! - \[FunctionGemma: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/functiongemma.md): Learn how to run and fine-tune FunctionGemma locally on your device and phone. - \[DeepSeek-V3.1: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/deepseek-v3.1-how-to-run-locally.md): A guide on how to run DeepSeek-V3.1 and Terminus on your own local device! - \[DeepSeek-R1-0528: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md): A guide on how to run DeepSeek-R1-0528 including Qwen3 on your own local device! - \[Liquid LFM2.5: How To Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/lfm2.5.md): Run and fine-tune LFM2.5 Instruct and Vision locally on your device! - \[Magistral: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/magistral-how-to-run-and-fine-tune.md): Meet Magistral - Mistral's new reasoning models. - \[IBM Granite 4.0\](https://unsloth.ai/docs/models/tutorials/ibm-granite-4.0.md): How to run IBM Granite-4.0 with Unsloth GGUFs on llama.cpp, Ollama and how to fine-tune! - \[Llama 4: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/llama-4-how-to-run-and-fine-tune.md): How to run Llama 4 locally using our dynamic GGUFs which recovers accuracy compared to standard quantization. - \[Grok 2\](https://unsloth.ai/docs/models/tutorials/grok-2.md): Run xAI's Grok 2 model locally! - \[Devstral: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/devstral-how-to-run-and-fine-tune.md): Run and fine-tune Mistral Devstral 1.1, including Small-2507 and 2505. - \[How to Run Local LLMs with Docker: Step-by-Step Guide\](https://unsloth.ai/docs/models/tutorials/how-to-run-llms-with-docker.md): Learn how to run Large Language Models (LLMs) with Docker & Unsloth on your local device. - \[DeepSeek-V3-0324: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/deepseek-v3-0324-how-to-run-locally.md): How to run DeepSeek-V3-0324 locally using our dynamic quants which recovers accuracy - \[DeepSeek-R1: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/deepseek-r1-how-to-run-locally.md): A guide on how you can run our 1.58-bit Dynamic Quants for DeepSeek-R1 using llama.cpp. - \[DeepSeek-R1 Dynamic 1.58-bit\](https://unsloth.ai/docs/models/tutorials/deepseek-r1-how-to-run-locally/deepseek-r1-dynamic-1.58-bit.md): See performance comparison tables for Unsloth's Dynamic GGUF Quants vs Standard IMatrix Quants. - \[Phi-4 Reasoning: How to Run & Fine-tune\](https://unsloth.ai/docs/models/tutorials/phi-4-reasoning-how-to-run-and-fine-tune.md): Learn to run & fine-tune Phi-4 reasoning models locally with Unsloth + our Dynamic 2.0 quants - \[QwQ-32B: How to Run effectively\](https://unsloth.ai/docs/models/tutorials/qwq-32b-how-to-run-effectively.md): How to run QwQ-32B effectively with our bug fixes and without endless generations + GGUFs. - \[Cogito v2.1: How to Run Locally\](https://unsloth.ai/docs/models/tutorials/cogito-v2-how-to-run-locally.md): Cogito v2.1 LLMs are one of the strongest open models in the world trained with IDA. Also v1 comes in 4 sizes: 70B, 109B, 405B and 671B, allowing you to select which size best matches your hardware. - \[GLM-5.2 - How to Run Locally\](https://unsloth.ai/docs/models/glm-5.2.md): Run the new GLM-5.2 model by Z.ai on local hardware! - \[Qwen3.6 - How to Run Locally\](https://unsloth.ai/docs/models/qwen3.6.md): Run the new Qwen3.6-27B and 35B-A3B models locally! - \[Gemma 4 - How to Run Locally\](https://unsloth.ai/docs/models/gemma-4.md): Run Google’s new Gemma 4 models locally, including E2B, E4B, 26B A4B, and 31B. - \[Gemma 4 QAT\](https://unsloth.ai/docs/models/gemma-4/qat.md): Run Google Gemma 4 QAT models locally, including E2B, E4B, 12B, 26B-A4B, and 31B. - \[Gemma 4 Fine-tuning Guide\](https://unsloth.ai/docs/models/gemma-4/train.md): Train Gemma 4 by Google with Unsloth. - \[DeepSeek-V4: How to Run Locally\](https://unsloth.ai/docs/models/deepseek-v4.md): Run DeepSeek-V4-Flash locally on your own device! - \[Inkling - How to Run Locally\](https://unsloth.ai/docs/models/inkling.md): Learn how to run Thinking Machine Labs' Inkling multimodal model locally. - \[Kimi K2.7 Code - How to Run Locally\](https://unsloth.ai/docs/models/kimi-k2.7-code.md): Step-by-step guide to running Kimi K2.7 Code on your own local device. - \[How to Run MTP Models: Multi-Token Prediction Guide\](https://unsloth.ai/docs/models/mtp.md) - \[DiffusionGemma - How to Run Locally\](https://unsloth.ai/docs/models/diffusiongemma.md) - \[MiniMax M3 - How to Run Locally\](https://unsloth.ai/docs/models/minimax-m3.md): Run MiniMax M3 LLM locally on your own device! - \[Qwen3.5 - How to Run Locally\](https://unsloth.ai/docs/models/qwen3.5.md): Run the new Qwen3.5 LLMs including Medium: Qwen3.5-35B-A3B, 27B, 122B-A10B, Small: Qwen3.5-0.8B, 2B, 4B, 9B and 397B-A17B on your local device! - \[Qwen3.5 Fine-tuning Guide\](https://unsloth.ai/docs/models/qwen3.5/fine-tune.md): Learn how to fine-tune Qwen3.5 LLMs with Unsloth. - \[Qwen3.5 GGUF Benchmarks\](https://unsloth.ai/docs/models/qwen3.5/gguf-benchmarks.md): See how Unsloth Dynamic GGUFs perform + analysis of perplexity, KL divergence & MXFP4. - \[Kimi K2.6 - How to Run Locally\](https://unsloth.ai/docs/models/kimi-k2.6.md): Step-by-step guide to running Kimi-K2.6 on your own local device. - \[NVIDIA Nemotron 3 Ultra - How To Run Locally\](https://unsloth.ai/docs/models/nemotron-3-ultra.md): Run Nemotron-3-Ultra-550B-A55B locally on your device! - \[NVIDIA Nemotron 3 Nano Omni - How To Run Locally\](https://unsloth.ai/docs/models/nemotron-3-nano-omni.md): Run & fine-tune Nemotron-3-Nano-Omni-30B-A3B locally on your device! - \[Mistral 3.5 - How To Run Locally\](https://unsloth.ai/docs/models/mistral-3.5.md): Guide for Mistral Mistral 3.5 models, to run or fine-tune locally on your device - \[IBM Granite 4.1 - How to Run Locally\](https://unsloth.ai/docs/models/ibm-granite-4.1.md): Run IBM Granite-4.1 with Unsloth GGUFs and how to fine-tune! - \[GLM-5.1 - How to Run Locally\](https://unsloth.ai/docs/models/glm-5.1.md): Run the new GLM-5.1 model by Z.ai on your own local device! - \[Qwen3-Coder-Next: How to Run Locally\](https://unsloth.ai/docs/models/qwen3-coder-next.md): Guide to run Qwen3-Coder-Next locally on your device! - \[NVIDIA Nemotron 3 Nano - How To Run Guide\](https://unsloth.ai/docs/models/nemotron-3.md): Run & fine-tune NVIDIA Nemotron 3 Nano locally on your device! - \[NVIDIA Nemotron-3-Super: How To Run Guide\](https://unsloth.ai/docs/models/nemotron-3/nemotron-3-super.md): Run & fine-tune NVIDIA Nemotron-3-Super-120B-A12B locally on your device! - \[gpt-oss: How to Run Guide\](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune.md): Run & fine-tune OpenAI's new open-source models! - \[gpt-oss Reinforcement Learning\](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune/gpt-oss-reinforcement-learning.md) - \[Tutorial: How to Train gpt-oss with RL\](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune/gpt-oss-reinforcement-learning/tutorial-how-to-train-gpt-oss-with-rl.md): Learn to train OpenAI gpt-oss with GRPO to autonomously beat 2048 locally or on Colab. - \[Tutorial: How to Fine-tune gpt-oss\](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune/tutorial-how-to-fine-tune-gpt-oss.md): Learn step-by-step how to train OpenAI gpt-oss locally with Unsloth. - \[Long Context gpt-oss Training\](https://unsloth.ai/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md) - \[How to use Unsloth as an API endpoint\](https://unsloth.ai/docs/basics/api.md) - \[Inference & Deployment\](https://unsloth.ai/docs/basics/inference-and-deployment.md): Learn how to save your finetuned model so you can run it in your favorite inference engine. - \[Saving to GGUF\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf.md) - \[Speculative Decoding\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf/speculative-decoding.md): Speculative Decoding with llama-server, llama.cpp, vLLM and more for 2x faster inference - \[vLLM Deployment & Inference Guide\](https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide.md): Guide on saving and deploying LLMs to vLLM for serving LLMs in production - \[vLLM Engine Arguments\](https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/vllm-engine-arguments.md) - \[LoRA Hot Swapping Guide\](https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/lora-hot-swapping-guide.md) - \[Saving models to Ollama\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-ollama.md) - \[Deploying models to LM Studio\](https://unsloth.ai/docs/basics/inference-and-deployment/lm-studio.md): Saving models to GGUF so you can run and deploy them to LM Studio - \[How to install LM Studio CLI in Linux Terminal\](https://unsloth.ai/docs/basics/inference-and-deployment/lm-studio/how-to-install-lm-studio-cli-in-linux-terminal.md): LM Studio CLI installation guide without a UI in a terminal instance. - \[SGLang Deployment & Inference Guide\](https://unsloth.ai/docs/basics/inference-and-deployment/sglang-guide.md): Guide on saving and deploying LLMs to SGLang for serving LLMs in production - \[Unsloth Inference\](https://unsloth.ai/docs/basics/inference-and-deployment/unsloth-inference.md): Learn how to run your finetuned model with Unsloth's faster inference. - \[llama-server & OpenAI endpoint Deployment Guide\](https://unsloth.ai/docs/basics/inference-and-deployment/llama-server-and-openai-endpoint.md): Deploying via llama-server with an OpenAI compatible endpoint - \[How to Run and Deploy LLMs on your iOS or Android Phone\](https://unsloth.ai/docs/basics/inference-and-deployment/deploy-llms-phone.md): Tutorial for fine-tuning your own LLM and deploying it on your Android or iPhone with ExecuTorch. - \[Troubleshooting Inference\](https://unsloth.ai/docs/basics/inference-and-deployment/troubleshooting-inference.md): If you're experiencing issues when running or saving your model. - \[Deploying LLMs with Hugging Face Jobs\](https://unsloth.ai/docs/basics/inference-and-deployment/deploying-llms-with-hugging-face-jobs.md): Using Hugging Face jobs and skills to fine-tune LFM with Codex / Claude Code with a SKILL. - \[How to Run Local LLMs with Claude Code\](https://unsloth.ai/docs/basics/claude-code.md): Guide to use open models with Claude Code on your local device. - \[How to Run Local LLMs with OpenAI Codex\](https://unsloth.ai/docs/basics/codex.md): Use open models with OpenAI Codex on your device locally. - \[Run Unsloth Dynamic NVFP4 Guide\](https://unsloth.ai/docs/basics/nvfp4.md): Learn how Unsloth Dynamic NVFP4 enables fast, accurate 4-bit inference on NVIDIA Blackwell GPUs. - \[Train & run models on AMD GPUs with Unsloth\](https://unsloth.ai/docs/basics/amd.md) - \[How to Use MCP Servers with Local LLMs\](https://unsloth.ai/docs/basics/mcp.md): Learn how to connect MCP Servers to open AI models with screenshots. - \[Multi-GPU Fine-tuning with Unsloth\](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth.md): Learn how to fine-tune LLMs on multiple GPUs and parallelism with Unsloth. - \[Multi-GPU Fine-tuning with Distributed Data Parallel (DDP)\](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp.md): Learn how to use the Unsloth CLI to train on multiple GPUs with Distributed Data Parallel (DDP)! - \[Fine-tuning Embedding Models with Unsloth Guide\](https://unsloth.ai/docs/basics/embedding-finetuning.md): Learn how to easily fine-tune embedding models with Unsloth. - \[Fine-tune MoE Models 12x Faster with Unsloth\](https://unsloth.ai/docs/basics/faster-moe.md): Train MoE LLMs locally using Unsloth Guide. - \[Text-to-Speech (TTS) Fine-tuning Guide\](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning.md): Learn how to to fine-tune TTS & STT voice models with Unsloth. - \[Unsloth Dynamic 2.0 GGUFs\](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs.md): A big new upgrade to our Dynamic Quants! - \[Unsloth Dynamic GGUFs on Aider Polyglot\](https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot.md): Performance of Unsloth Dynamic GGUFs on Aider Polyglot Benchmarks - \[Tool Calling Guide for Local LLMs\](https://unsloth.ai/docs/basics/tool-calling-guide-for-local-llms.md) - \[Vision Fine-tuning\](https://unsloth.ai/docs/basics/vision-fine-tuning.md): Learn how to fine-tune vision/multimodal LLMs with Unsloth - \[Troubleshooting & FAQs\](https://unsloth.ai/docs/basics/troubleshooting-and-faqs.md): Tips to solve issues, and frequently asked questions. - \[Hugging Face Hub, XET debugging\](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md): Debugging, troubleshooting stalled, stuck downloads and slow downloads - \[Chat Templates\](https://unsloth.ai/docs/basics/chat-templates.md): Learn the fundamentals and customization options of chat templates, including Conversational, ChatML, ShareGPT, Alpaca formats, and more! - \[Unsloth Environment Flags\](https://unsloth.ai/docs/basics/unsloth-environment-flags.md): Advanced flags which might be useful if you see breaking finetunes, or you want to turn stuff off. - \[Continued Pretraining\](https://unsloth.ai/docs/basics/continued-pretraining.md): AKA as Continued Finetuning. Unsloth allows you to continually pretrain so a model can learn a new language. - \[Finetuning from Last Checkpoint\](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint.md): Checkpointing allows you to save your finetuning progress so you can pause it and then continue. - \[Unsloth Benchmarks\](https://unsloth.ai/docs/basics/unsloth-benchmarks.md): Unsloth recorded benchmarks on NVIDIA GPUs. - \[Run Coding Agents with Local LLMs using Unsloth Start\](https://unsloth.ai/docs/integrations/unsloth-start.md) - \[Connect API Providers & Model Servers to Unsloth\](https://unsloth.ai/docs/integrations/connections.md): Guide to connect OpenAI, Anthropic, Ollama, llama.cpp, vLLM and other providers to Unsloth. Add API keys or model server URLs, load models, and use external models in chat. - \[Connect OpenAI to Unsloth: Run GPT Models in Local Chat\](https://unsloth.ai/docs/integrations/connections/openai.md) - \[Connect Anthropic to Unsloth: Run Claude Models in Local Chat\](https://unsloth.ai/docs/integrations/connections/anthropic-claude.md) - \[Connect llama.cpp to Unsloth: Run GGUFs with llama-server\](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md) - \[Connect vLLM to Unsloth for Local Chat Inference\](https://unsloth.ai/docs/integrations/connections/vllm.md) - \[How to Connect Ollama to Unsloth\](https://unsloth.ai/docs/integrations/connections/ollama.md) - \[How to Connect OpenRouter to Unsloth: API Key & Model Setup\](https://unsloth.ai/docs/integrations/connections/openrouter.md) - \[How to Run Local AI Models with Hermes Agent\](https://unsloth.ai/docs/integrations/hermes-agent.md): Guide on using open LLMs with Hermes Agent locally. - \[How to Run Local AI Models with OpenClaw\](https://unsloth.ai/docs/integrations/openclaw.md): Guide to running local LLMs with OpenClaw. - \[How to Run Local AI Models with OpenCode\](https://unsloth.ai/docs/integrations/opencode.md): Guide to connect open LLMs with OpenCode on your local device. - \[Connect Python SDK to Unsloth\](https://unsloth.ai/docs/integrations/connect-python-sdk-to-unsloth.md): Guide to calling Unsloth's local API from Python using the official OpenAI or Anthropic SDKs including streaming, vision, function calling, and Unsloth's built-in server-side tools. - \[Connect Curl & HTTP to Unsloth\](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth.md): Guide to hitting Unsloth's API with curl (or any HTTP client), complete with copy-pasteable recipes for every endpoint and feature.. - \[3x Faster LLM Training with Unsloth Kernels + Packing\](https://unsloth.ai/docs/blog/3x-faster-training-packing.md): Learn how Unsloth increases training throughput and eliminates padding waste for fine-tuning. - \[500K Context Length Fine-tuning\](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning.md): Learn how to enable >500K token context window fine-tuning with Unsloth. - \[Quantization-Aware Training (QAT)\](https://unsloth.ai/docs/blog/quantization-aware-training-qat.md): Quantize models to 4-bit with Unsloth and PyTorch to recover accuracy. - \[Fine-Tuning LLMs on NVIDIA DGX Station with Unsloth\](https://unsloth.ai/docs/blog/dgx-station.md): NVIDIA DGX Station tutorial on how to fine-tune with notebooks from Unsloth. - \[How to Fine-tune LLMs with Unsloth & Docker\](https://unsloth.ai/docs/blog/how-to-fine-tune-llms-with-unsloth-and-docker.md): Learn how to fine-tune LLMs or do Reinforcement Learning (RL) with Unsloth's Docker image. - \[Fine-tuning LLMs with NVIDIA DGX Spark and Unsloth\](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth.md): Tutorial on how to fine-tune and do reinforcement learning (RL) with OpenAI gpt-oss on NVIDIA DGX Spark. - \[Fine-tuning LLMs with Blackwell, RTX 50 series & Unsloth\](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md): Learn how to fine-tune LLMs on NVIDIA's Blackwell RTX 50 series and B200 GPUs with our step-by-step guide. - \[How to Run Diffusion Image GGUFs in ComfyUI\](https://unsloth.ai/docs/blog/comfyui.md): Guide for running Unsloth Diffusion GGUF models in ComfyUI. - \[AI Engineer's 2025\](https://unsloth.ai/docs/blog/ai-engineers-2025.md): Slides to our AI Engineer's Worlds Fair 2025 Workshop. - \[Unsloth - AI Engineer World's Fair 2026\](https://unsloth.ai/docs/blog/ai-engineer-2026.md): Slides to our AI Engineer's Worlds Fair 2026 Workshop. - \[GPU Mode - Reinforcement Learning Mini Conference 2026\](https://unsloth.ai/docs/blog/gpu-mode-conference.md): Slides to our PyTorch Conference 2025 Talks. - \[Unsloth AMD PyTorch Synthetic Data Hackathon\](https://unsloth.ai/docs/blog/unsloth-amd-pytorch-synthetic-data-hackathon.md): Tips & tricks, troubleshooting and guide to run Unsloth on an AMD GPU. - \[PyTorch Conference 2025 - Unsloth\](https://unsloth.ai/docs/blog/pytorch-conference-2025-unsloth.md): Slides to our PyTorch Conference 2025 Talks. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on a page URL with the \`ask\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/readme.md?ask= \`\`\` The question should be specific, self-contained, and written in natural language. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/amd.md). # Fine-tuning LLMs on AMD GPUs with Unsloth Guide Fine-tune LLMs up to 2x faster with \\~70% less memory on AMD hardware, no NVIDIA required. Unsloth supports AMD Radeon RDNA 3/3.5/4 (RX 6000–9000 series) on both Windows and Linux as well as data center GPUs including the MI300X (192GB). {% stepper %} {% step %} #### \*\*One-line installer\*\* \*\*Easiest install:\*\* Skip all the steps below with the one-line installer, it auto-detects your AMD GPU, installs ROCm-optimized PyTorch, bitsandbytes, and launches Unsloth Studio: \*\*Linux :\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows (PowerShell):\*\* \`\`\`powershell irm https://unsloth.ai/install.ps1 | iex \`\`\` The manual steps below are for users who want to install the Unsloth Python library on AMD with the required dependencies. {% endstep %} {% step %} #### \*\*Make a new isolated environment (Optional)\*\* To not break any system packages, you can make an isolated pip environment. Reminder to check what Python version you have! It might be \`pip3\`, \`pip3.13\`, \`python3\`, \`python.3.13\` etc. \*\*Linux:\*\* 3.13 shown; any 3.11-3.13 works everywhere (3.10 works for manual installs {% code overflow="wrap" %} \`\`\`bash # Linux — swap 3.13 for whichever 3.10-3.13 you have apt update && apt install python3.13-venv -y python3.13 -m venv unsloth\_env source unsloth\_env/bin/activate pip install uv \`\`\` {% endcode %} \*\*Windows (PowerShell):\*\* 3.13 shown; use 3.12 if installing the unsloth\\\[rocm72-torch291\] extra \`\`\`shellscript py -3.13 -m venv unsloth\_env unsloth\_env\\Scripts\\Activate.ps1 pip install uv \`\`\` {% endstep %} {% step %} #### \*\*Install PyTorch\*\* Planning to install Unsloth with an AMD extra (\`unsloth\[rocm72-torch291\]\` etc., see the next section)? Those bundle a matching PyTorch, so you can \*\*skip this section\*\*. Install PyTorch here only if you want to manage it yourself or your ROCm version isn't covered by an extra. \*\*Linux:\*\* \\ Install PyTorch with ROCm support from the PyTorch index. Check your ROCm version via \`amd-smi version\` (look for the \`ROCm version:\` line), then change \`https://download.pytorch.org/whl/rocm7.1\` to match it. ROCm 6.0 or newer is required. {% code overflow="wrap" %} \`\`\`bash uv pip install "torch>=2.4,<2.11.0" "torchvision<0.26.0" "torchaudio<2.11.0" \\ --index-url https://download.pytorch.org/whl/rocm7.1 --upgrade --force-reinstall \`\`\` {% endcode %} ROCm 7.2 ships newer wheels (torch 2.11), so on ROCm 7.2 use this instead: \`\`\`bash uv pip install "torch>=2.11.0,<2.12.0" torchvision torchaudio \\ --index-url https://download.pytorch.org/whl/rocm7.2 --upgrade --force-reinstall \`\`\` Available index tags are \`rocm6.0\`, \`rocm6.1\`, \`rocm6.2\`, \`rocm6.3\`, \`rocm6.4\`, \`rocm7.0\`, \`rocm7.1\`, and \`rocm7.2\`. ROCm 6.5-6.9 has no dedicated wheels (use \`rocm6.4\`), and ROCm 7.3+ uses \`rocm7.2\`. \*The version caps prevent accidentally pulling torch 2.11+ which only has ROCm 7.2 wheels and will break things. Update \`rocm7.1\` to match your detected version as before.\* We also wrote a single terminal command to extract the correct ROCM version if it helps. \`\`\`bash ROCM\_TAG="$({ command -v amd-smi >/dev/null 2>&1 && amd-smi version 2>/dev/null | awk -F'ROCm version: ' 'NF>1{split($2,a,"."); print "rocm"a\[1\]"."a\[2\]; ok=1; exit} END{exit !ok}'; } || { \[ -r /opt/rocm/.info/version \] && awk -F. '{print "rocm"$1"."$2; exit}' /opt/rocm/.info/version; } || { command -v hipconfig >/dev/null 2>&1 && hipconfig --version 2>/dev/null | awk -F': \*' '/HIP version/{split($2,a,"."); print "rocm"a\[1\]"."a\[2\]; ok=1; exit} END{exit !ok}'; } || { command -v dpkg-query >/dev/null 2>&1 && ver="$(dpkg-query -W -f="${Version}\\n" rocm-core 2>/dev/null)" && \[ -n "$ver" \] && awk -F'\[.-\]' '{print "rocm"$1"."$2; exit}' <<<"$ver"; } || { command -v rpm >/dev/null 2>&1 && ver="$(rpm -q --qf '%{VERSION}\\n' rocm-core 2>/dev/null)" && \[ -n "$ver" \] && awk -F'\[.-\]' '{print "rocm"$1"."$2; exit}' <<<"$ver"; })"; \[ -n "$ROCM\_TAG" \] && uv pip install "torch>=2.4,<2.11.0" "torchvision<0.26.0" "torchaudio<2.11.0" --index-url "https://download.pytorch.org/whl/$ROCM\_TAG" --upgrade --force-reinstall \`\`\` \*Note: If your ROCm version is 7.2 or higher, replace \`$ROCM\_TAG\` in the command above with \`rocm7.1,\` no PyTorch wheels exist yet for 7.2+.\* ![](https://unsloth.ai/files/9lCfcDCiR5HW4tTepAH2) {% endstep %} {% step %} #### \*\*Install Unsloth\*\* Install Unsloth with AMD extras: {% code overflow="wrap" %} \`\`\`bash uv pip install unsloth\[amd\] \`\`\` {% endcode %} ![](https://unsloth.ai/files/MlkhF0QUVdc2rjJBEytX) ⚠️ Required for AMD: install ROCm-compatible bitsandbytes\\ All ROCm systems need a pre-release bitsandbytes build, versions ≤ 0.49.2 have a 4-bit decode NaN bug on every AMD GPU. Note: use \`pip\` not \`uv\` for this step, \`uv\` rejects the pre-release wheel due to a version mismatch in the filename. {% code overflow="wrap" %} \`\`\`bash # x86\_64 systems: pip install --force-reinstall --no-cache-dir --no-deps \\ "https://github.com/bitsandbytes-foundation/bitsandbytes/releases/download/continuous-release\_main/bitsandbytes-1.33.7.preview-py3-none-manylinux\_2\_24\_x86\_64.whl" # aarch64 systems: replace x86\_64 with aarch64 in the URL above # Fallback if the URL is unreachable: # pip install --force-reinstall --no-cache-dir --no-deps "bitsandbytes>=0.49.1" \`\`\` {% endcode %} ![](https://unsloth.ai/files/XN8Z2xoE9pGXLttLiwrz) {% endstep %} {% step %} #### \*\*Start fine-tuning with Unsloth!\*\* And that's it. Try some examples in our \[\*\*Unsloth Notebooks\*\*\](/docs/get-started/unsloth-notebooks.md) page! You can view our dedicated \[fine-tuning\](/docs/get-started/fine-tuning-llms-guide.md) or \[reinforcement learning\](/docs/get-started/reinforcement-learning-rl-guide.md) guides. Heres a brief example as well: \*\*1. Set environment variables\*\* {% code overflow="wrap" %} \`\`\`bash export HSA\_OVERRIDE\_GFX\_VERSION=9.4.2 # Required for AMD MI300X export HF\_HUB\_DISABLE\_XET=1 # Fixes HuggingFace download issues on AMD \`\`\` {% endcode %} \*\*\*Note:\*\*\* \*\`HSA\_OVERRIDE\_GFX\_VERSION=9.4.2\` tells ROCm to treat your GPU as gfx942 (MI300X). Without this, some kernels may fail to compile or run.\* \*\*2. Load and configure model\*\* {% code overflow="wrap" %} \`\`\`python from unsloth import FastModel model, tokenizer = FastModel.from\_pretrained( model\_name = "unsloth/gemma-4-26b-a4b-it", max\_seq\_length = 2048, load\_in\_4bit = True, ) model = FastModel.get\_peft\_model( model, r = 16, lora\_alpha = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj"\], ) \`\`\` {% endcode %} \*\*3. Train\*\* {% code overflow="wrap" %} \`\`\`python from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( model = model, tokenizer = tokenizer, train\_dataset = dataset, formatting\_func = formatting\_func, args = SFTConfig( per\_device\_train\_batch\_size = 1, gradient\_accumulation\_steps = 4, max\_steps = 60, output\_dir = "outputs", report\_to = "none", ), ) trainer\_stats = trainer.train() \`\`\` {% endcode %} ![](https://unsloth.ai/files/f2ZKFGDdYUKCuHU5F3SS) \*\*\*Note:\*\* On AMD GPUs, Flash Attention 2 is not available. Unsloth automatically falls back to Xformers, which provides equivalent performance on ROCm. The warning can be safely ignored.\* {% endstep %} {% endstepper %} ### :1234: Reinforcement Learning on AMD GPUs You can use our :ledger:\[gpt-oss RL auto win 2048\](https://github.com/unslothai/notebooks/blob/main/nb/AMD-gpt\_oss\_\\(20B\\)\_Reinforcement\_Learning\_2048\_Game\_BF16.ipynb) example on a MI300X (192GB) GPU. The goal is to play the 2048 game automatically and win it with RL. The LLM (gpt-oss 20b) auto devises a strategy to win the 2048 game, and we calculate a high reward for winning strategies, and low rewards for failing strategies. {% columns %} {% column %} ![](https://unsloth.ai/files/SSEOB6ZRI1FyoralepNg) {% endcolumn %} {% column %} The reward over time is increasing after around 300 steps or so! The goal for RL is to maximize the average reward to win the 2048 game. ![](https://unsloth.ai/files/Z1iVqHe8eHfWSdc4j2sI) {% endcolumn %} {% endcolumns %} We used an AMD MI300X machine (192GB) to run the 2048 RL example with Unsloth, and it worked well! ![](https://unsloth.ai/files/NDfxHBQ4X6Qw1LnZFiFO) ![](https://unsloth.ai/files/PZ4STHvi1o5M4mI7nAlH) You can also use our :ledger:\[automatic kernel gen RL notebook\](https://github.com/unslothai/notebooks/blob/main/nb/AMD-gpt\_oss\_\\(20B\\)\_GRPO\_BF16.ipynb) also with gpt-oss to auto create matrix multiplication kernels in Python. The notebook also devices multiple methods to counteract reward hacking. {% columns %} {% column width="50%" %} The prompt we used to auto create these kernels was: {% code overflow="wrap" %} \`\`\`\` Create a new fast matrix multiplication function using only native Python code. You are given a list of list of numbers. Output your new function in backticks using the format below: \`\`\` python def matmul(A, B): return ... \`\`\` \`\`\`\` {% endcode %} {% endcolumn %} {% column width="50%" %} The RL process learns for example how to apply the Strassen algorithm for faster matrix multiplication inside of Python. ![](https://unsloth.ai/files/fReUWthRUkDyj6srHdG1) {% endcolumn %} {% endcolumns %} ### :books:AMD Free One-click notebooks AMD provides one-click notebooks equipped with \*\*free 192GB VRAM MI300X GPUs\*\* through their Dev Cloud. Train large models completely for free (no signup or credit card required): \* \[Qwen3 (32B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(32B\\)\_A100-Reasoning-Conversational.ipynb) \* \[Llama 3.3 (70B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Llama3.3\_\\(70B\\)\_A100-Conversational.ipynb) \* \[Qwen3 (14B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* \[Mistral v0.3 (7B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-Mistral\_v0.3\_\\(7B\\)-Alpaca.ipynb) \* \[GPT OSS MXFP4 (20B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/AMD-GPT\_OSS\_MXFP4\_\\(20B\\)-Inference.ipynb) - Inference \* \[Gemma4 (E2B)\](https://amd-ai-academy.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) - RL Sudoku \* Unsloth Studio {% embed url="" %} You can use any Unsloth notebook by prepending in \[Unsloth Notebooks\](/docs/get-started/unsloth-notebooks.md) by changing the link from \\ to {% columns %} {% column width="33.33333333333333%" %} ![](https://unsloth.ai/files/l2ybUSNrZQlRBFB21k5N) {% endcolumn %} {% column width="66.66666666666667%" %} ![](https://unsloth.ai/files/ksXwklNfCmnTYRPEuPBA) {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/amd.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md). # DeepSeek-OCR: How to Run & Fine-tune \*\*DeepSeek-OCR\*\* is a 3B-parameter vision model for OCR and document understanding. It uses \*context optical compression\* to convert 2D layouts into vision tokens, enabling efficient long-context processing. Capable of handling tables, papers, and handwriting, DeepSeek-OCR achieves 97% precision while using 10× fewer vision tokens than text tokens - making it 10× more efficient than text-based LLMs. You can fine-tune DeepSeek-OCR to enhance its vision or language performance. In our Unsloth \[\*\*free fine-tuning notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb), we demonstrated a \[88.26% improvement\](#fine-tuning-deepseek-ocr) for language understanding. [Running DeepSeek-OCR](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md#running-deepseek-ocr) [Fine-tuning DeepSeek-OCR](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md#fine-tuning-deepseek-ocr) > \*\*Our model upload that enables fine-tuning + more inference support:\*\* \[\*\*DeepSeek-OCR\*\*\](https://huggingface.co/unsloth/DeepSeek-OCR) ## 🖥️ \*\*Running DeepSeek-OCR\*\* To run the model in \[vLLM\](#vllm-run-deepseek-ocr-tutorial) or \[Unsloth\](#unsloth-run-deepseek-ocr-tutorial), here are the recommended settings: ### :gear: Recommended Settings DeepSeek recommends these settings: \* \*\*Temperature = 0.0\*\* \* \`max\_tokens = 8192\` \* \`ngram\_size = 30\` \* \`window\_size = 90\` ### 📖 vLLM: Run DeepSeek-OCR Tutorial 1. Obtain the latest \`vLLM\` via: \`\`\`bash uv venv source .venv/bin/activate # Until v0.11.1 release, you need to install vLLM from nightly build uv pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightly \`\`\` 2. Then run the following code: {% code overflow="wrap" %} \`\`\`python from vllm import LLM, SamplingParams from vllm.model\_executor.models.deepseek\_ocr import NGramPerReqLogitsProcessor from PIL import Image # Create model instance llm = LLM( model="unsloth/DeepSeek-OCR", enable\_prefix\_caching=False, mm\_processor\_cache\_gb=0, logits\_processors=\[NGramPerReqLogitsProcessor\] ) # Prepare batched input with your image file image\_1 = Image.open("path/to/your/image\_1.png").convert("RGB") image\_2 = Image.open("path/to/your/image\_2.png").convert("RGB") prompt = "\\nFree OCR." model\_input = \[ { "prompt": prompt, "multi\_modal\_data": {"image": image\_1} }, { "prompt": prompt, "multi\_modal\_data": {"image": image\_2} } \] sampling\_param = SamplingParams( temperature=0.0, max\_tokens=8192, # ngram logit processor args extra\_args=dict( ngram\_size=30, window\_size=90, whitelist\_token\_ids={128821, 128822}, # whitelist: , ), skip\_special\_tokens=False, ) # Generate output model\_outputs = llm.generate(model\_input, sampling\_param) # Print output for output in model\_outputs: print(output.outputs\[0\].text) \`\`\` {% endcode %} ### 🦥 Unsloth: Run DeepSeek-OCR Tutorial 1. Obtain the latest \`unsloth\` via \`pip install --upgrade unsloth\` . If you already have Unsloth, update it via \`pip install --upgrade --force-reinstall --no-deps --no-cache-dir unsloth unsloth\_zoo\` 2. Then use the code below to run DeepSeek-OCR: {% code overflow="wrap" %} \`\`\`python from unsloth import FastVisionModel import torch from transformers import AutoModel import os os.environ\["UNSLOTH\_WARN\_UNINITIALIZED"\] = '0' from huggingface\_hub import snapshot\_download snapshot\_download("unsloth/DeepSeek-OCR", local\_dir = "deepseek\_ocr") model, tokenizer = FastVisionModel.from\_pretrained( "./deepseek\_ocr", load\_in\_4bit = False, # Use 4bit to reduce memory use. False for 16bit LoRA. auto\_model = AutoModel, trust\_remote\_code = True, unsloth\_force\_compile = True, use\_gradient\_checkpointing = "unsloth", # True or "unsloth" for long context ) prompt = "\\nFree OCR. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 640, crop\_mode=True, save\_results = True, test\_compress = False) \`\`\` {% endcode %} ## 🦥 \*\*Fine-tuning DeepSeek-OCR\*\* Unsloth supports fine-tuning of DeepSeek-OCR. Since the default model isn't runnable on the latest \`transformers\` version, we added changes from the \[Stranger Vision HF\](https://huggingface.co/strangervisionhf) team, to then enable inference. As usual, Unsloth trains DeepSeek-OCR 1.4x faster with 40% less VRAM and 5x longer context lengths - no accuracy degradation.\\ \\ We created two free DeepSeek-OCR Colab notebooks (with and without eval): \* DeepSeek-OCR: \[Fine-tuning only notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb) \* DeepSeek-OCR: \[Fine-tuning + Evaluation notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\)-Eval.ipynb) (A100) Fine-tuning DeepSeek-OCR on a 200K sample Persian dataset resulted in substantial gains in Persian text detection and understanding. We evaluated the base model against our fine-tuned version on 200 Persian transcript samples, observing an \*\*88.26% absolute improvement\*\* in Character Error Rate (CER). After only 60 training steps (batch size = 8), the mean CER decreased from \*\*149.07%\*\* to a mean of \*\*60.81%\*\*. This means the fine-tuned model is \*\*57%\*\* more accurate at understanding Persian. You can replace the Persian dataset with your own to improve DeepSeek-OCR for other use-cases.\\ \\ For replica-table eval results, use our eval notebook above. For detailed eval results, see below: ### Fine-tuned Evaluation Results: {% columns fullWidth="true" %} {% column %} \*\*DeepSeek-OCR Baseline\*\* Mean Baseline Model Performance: 149.07% CER for this eval set! \`\`\` ============================================================ Baseline Model Performance ============================================================ Number of samples: 200 Mean CER: 149.07% Median CER: 80.00% Std Dev: 310.39% Min CER: 0.00% Max CER: 3500.00% ============================================================ Best Predictions (Lowest CER): Sample 5024 (CER: 0.00%) Reference: چون هستی خیلی زیاد... Prediction: چون هستی خیلی زیاد... Sample 3517 (CER: 0.00%) Reference: تو ایران هیچوقت از اینها وجود نخواهد داشت... Prediction: تو ایران هیچوقت از اینها وجود نخواهد داشت... Sample 9949 (CER: 0.00%) Reference: کاش میدونستم هیچی بیخیال... Prediction: کاش میدونستم هیچی بیخیال... Worst Predictions (Highest CER): Sample 11155 (CER: 3500.00%) Reference: خسو... Prediction: \\\[ \\text{CH}\_3\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}... Sample 13366 (CER: 1900.00%) Reference: مشو... Prediction: \\\[\\begin{align\*}\\underline{\\mathfrak{su}}\_0\\end{align\*}\\\]... Sample 10552 (CER: 1014.29%) Reference: هیییییچ... Prediction: e \`\`\` {% endcolumn %} {% column %} \*\*DeepSeek-OCR Fine-tuned\*\* With 60 steps, we reduced CER from 149.07% to 60.43% (89% CER improvement)\ \ ============================================================\ Fine-tuned Model Performance\ ============================================================\ Number of samples: 200\ Mean CER: 60.43%\ Median CER: 50.00%\ Std Dev: 80.63%\ Min CER: 0.00%\ Max CER: 916.67%\ ============================================================\ \ Best Predictions (Lowest CER):\ \ Sample 301 (CER: 0.00%)\ Reference: باشه بابا تو لاکچری، تو خاص، تو خفن...\ Prediction: باشه بابا تو لاکچری، تو خاص، تو خفن...\ \ Sample 2512 (CER: 0.00%)\ Reference: از شخص حاج عبدالله زنجبیلی میگیرنش...\ Prediction: از شخص حاج عبدالله زنجبیلی میگیرنش...\ \ Sample 2713 (CER: 0.00%)\ Reference: نمی دونم والا تحمل نقد ندارن ظاهرا...\ Prediction: نمی دونم والا تحمل نقد ندارن ظاهرا...\ \ Worst Predictions (Highest CER):\ \ Sample 14270 (CER: 916.67%)\ Reference: ۴۳۵۹۴۷۴۷۳۸۹۰...\ Prediction: پروپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپیپریپریپریپریپریپریپریپریپریپریپریپریپریپر...\ \ Sample 3919 (CER: 380.00%)\ Reference: ۷۵۵۰۷۱۰۶۵۹...\ Prediction: وادووووووووووووووووووووووووووووووووووو...\ \ Sample 3718 (CER: 333.33%)\ Reference: ۳۲۶۷۲۲۶۵۵۸۴۶...\ Prediction: پُپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُ...\ \ \ {% endcolumn %} {% endcolumns %} An example from the 200K Persian dataset we used (you may use your own), showing the image on the left and the corresponding text on the right.\ \ ![](https://unsloth.ai/files/mPX1h636KA0l4JXyLnzf)\ \ \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks.md). # Notebooks Unsloth Entraînez votre propre modèle avec nos notebooks, alimentés par des calculs GPU gratuits. Cliquez sur Exécuter tout (ou enregistrez localement), ajoutez votre jeu de données, entraînez et déployez. Vous pouvez utiliser n’importe quel modèle dans les notebooks. [GRPO (RL)](https://unsloth.ai/pages/2bdfa5349e5d636154595876b11e5db5e1f9e9d6#grpo-reasoning-rl) [Synthèse vocale](https://unsloth.ai/pages/2bdfa5349e5d636154595876b11e5db5e1f9e9d6#text-to-speech-tts) [Vision](https://unsloth.ai/pages/2bdfa5349e5d636154595876b11e5db5e1f9e9d6#vision-multimodal) [Embedding](https://unsloth.ai/pages/2bdfa5349e5d636154595876b11e5db5e1f9e9d6#embedding-models) [Kaggle](https://unsloth.ai/pages/2bdfa5349e5d636154595876b11e5db5e1f9e9d6#kaggle-notebooks) Consultez aussi notre dépôt GitHub pour nos notebooks : \[github.com/unslothai/notebooks\](https://github.com/unslothai/notebooks/) ## Notebooks Colab \*\*Présentation de notre\*\* \[\*\*Unsloth Studio\*\*\](/docs/fr/nouveau/studio.md)✨ \*\*notebook.\*\* Entraînez et exécutez des modèles de moins de 22 milliards de paramètres : {% embed url="" %} ### Notebooks SFT standard : \* \[\*\*Gemma 4\*\*\](/docs/fr/modeles/gemma-4/train.md)\*\*:\*\* \[E4B \*\*(Vision)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E4B\\)-Vision.ipynb) \*\*•\*\* \[E2B \*\*(Texte)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)-Text.ipynb) \*\*•\*\* \[E2B \*\*(Audio)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)-Audio.ipynb) \*\*•\*\* \[\*\*31B\*\* (Kaggle)\](https://www.kaggle.com/code/danielhanchen/gemma4-31b-unsloth) \*\*•\*\* \[\*\*Inférence\*\*\](https://colab.research.google.com/github/unslothai/unsloth/blob/main/studio/Unsloth\_Studio\_Colab.ipynb) \* \[\*\*Qwen3.5\*\*\](/docs/fr/modeles/qwen3.5/fine-tune.md)\*\*:\*\* \[\*\*0,8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) • \[\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) • \[\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) \* \[gpt-oss (20b)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-Fine-tuning.ipynb) • \[Inférence\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/GPT\_OSS\_MXFP4\_\\(20B\\)-Inference.ipynb) • \[Fine-tuning\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-Fine-tuning.ipynb) \* \[EmbeddingGemma (300M)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/EmbeddingGemma\_\\(300M\\).ipynb) \* \[Qwen3 (14B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) • \[\*\*Qwen3-VL (8B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision.ipynb) \* \[\*\*Qwen3-2507-4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)-Instruct.ipynb) • \[Réflexion\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)-Thinking.ipynb) • \[Instruct\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)-Instruct.ipynb) \* \[Gemma 3 (4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) • \[Texte\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) • \[Vision\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) • \[270M\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(270M\\).ipynb) • \[\*\*FunctionGemma\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\).ipynb) \* \[Gemma 3n (E4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Conversational.ipynb) • \[Texte\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Conversational.ipynb) • \[Vision\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Vision.ipynb) • \[Audio\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Audio.ipynb) \* \[\*\*Mistral Ministral 3\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_VL\_\\(3B\\)\_Vision.ipynb) \* \[\*\*DeepSeek-OCR 2\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_2\_\\(3B\\).ipynb) \* \[IBM Granite-4.0-H\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Granite4.0.ipynb) \* \[Phi-4 (14B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* \[Llama 3.1 (8B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1\_\\(8B\\)-Alpaca.ipynb) • \[Llama 3.2 (1B + 3B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) ### GRPO (RL de raisonnement) : \* \[\*\*Gemma 4 E2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) - nouveau \* \[\*\*Qwen3.5 (4B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision\_GRPO.ipynb) - GRPO Vision \* \[gpt-oss-20b\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) (création automatique de kernels) \* \[Mistral Ministral 3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) (résolution de sudoku) - nouveau \* \[Qwen3-8B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_8B\_FP8\_GRPO.ipynb) (L4) - nouveau \* \[Llama-3.2-1B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama\_FP8\_GRPO.ipynb) (L4) - nouveau \* \[gpt-oss-20b\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_Reinforcement\_Learning\_2048\_Game.ipynb) (gagner automatiquement au jeu 2048) \* \[Qwen3-VL (8B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) - GSPO Vision \* \[Qwen3 (4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)-GRPO.ipynb) - LoRA GRPO avancé \* \[Gemma 3 (4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision-GRPO.ipynb) - GSPO Vision \* \[gpt-oss-20b\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/OpenEnv\_gpt\_oss\_\\(20B\\)\_Reinforcement\_Learning\_2048\_Game.ipynb) (exemple OpenEnv 2048) \* \[DeepSeek-R1-0528-Qwen3 (8B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/DeepSeek\_R1\_0528\_Qwen3\_\\(8B\\)\_GRPO.ipynb) (pour un cas d'utilisation multilingue) \* \[Gemma 3 (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(1B\\)-GRPO.ipynb) \* \[Llama 3.2 (3B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Advanced\_Llama3\_2\_\\(3B\\)\_GRPO\_LoRA.ipynb) - LoRA GRPO avancé \* \[Llama 3.1 (8B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1\_\\(8B\\)-GRPO.ipynb) \* \[Phi-4 (14B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4\_\\(14B\\)-GRPO.ipynb) \* \[Mistral v0.3 (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-GRPO.ipynb) \* \[Environnement multi-agents NeMo Gym \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/NeMo-Gym-Multi-Environment.ipynb)(Plusieurs environnements agentiques) ### Synthèse vocale (TTS) : \* \[Sesame-CSM (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Sesame\_CSM\_\\(1B\\)-TTS.ipynb) \* \[Orpheus-TTS (3B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Orpheus\_\\(3B\\)-TTS.ipynb) \* \[Whisper Large V3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Whisper.ipynb) - reconnaissance vocale (STT) \* \[Llasa-TTS (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llasa\_TTS\_\\(1B\\).ipynb) \* \[Spark-TTS (0.5B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Spark\_TTS\_\\(0\_5B\\).ipynb) \* \[Oute-TTS (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Oute\_TTS\_\\(1B\\).ipynb) \*\*Reconnaissance vocale (SST) :\*\* \* \[\*\*Gemma 4 (E2B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)-Audio.ipynb) \*\*- Audio - nouveau\*\* \* \[Whisper-Large-V3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Whisper.ipynb) \* \[Gemma 3n (E4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Audio.ipynb) - Audio ### Vision (multimodale) : \* \[\*\*Gemma 4\*\*\](/docs/fr/modeles/gemma-4/train.md)\*\*:\*\* \[E2B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(E2B\\)-Vision.ipynb) \*\*•\*\* \[\*\*31B\*\* (Kaggle)\](https://www.kaggle.com/code/danielhanchen/gemma4-31b-unsloth) - nouveau \* \[\*\*Qwen3.5\*\*\](/docs/fr/modeles/qwen3.5/fine-tune.md)\*\*:\*\* \[\*\*0,8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) • \[\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) • \[\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) - nouveau \* \[\*\*Qwen3-VL (8B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision.ipynb) \* \[\*\*Mistral Ministral 3\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_VL\_\\(3B\\)\_Vision.ipynb) \* \[\*\*DeepSeek-OCR\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb) \* \[\*\*Paddle-OCR (1B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Paddle\_OCR\_\\(1B\\)\_Vision.ipynb) \* \[Gemma 3n (E4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Vision.ipynb) \* \[Gemma 3 (4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) \* \[Llama 3.2 Vision (11B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(11B\\)-Vision.ipynb) \* \[Qwen2.5-VL (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5\_VL\_\\(7B\\)-Vision.ipynb) \* \[Pixtral (12B) 2409\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Pixtral\_\\(12B\\)-Vision.ipynb) \* \[Qwen3-VL\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) - GSPO Vision - nouveau \* \[Qwen2.5-VL\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2\_5\_7B\_VL\_GRPO.ipynb) - GSPO Vision \* \[Gemma 3 (4B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision-GRPO.ipynb) - GSPO Vision ### Modèles d'embedding : \* \[EmbeddingGemma (300M)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/EmbeddingGemma\_\\(300M\\).ipynb) - nouveau \* \[Qwen3-Embedding 4B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_Embedding\_\\(4B\\).ipynb) - nouveau \* \[Qwen3-Embedding 0.6B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_Embedding\_\\(0\_6B\\).ipynb) - nouveau \* \[BGE M3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/BGE\_M3.ipynb) - nouveau \* \[ModernBERT-large\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert\_classification.ipynb) - nouveau \* \[All-MiniLM-L6-v2\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/All\_MiniLM\_L6\_v2.ipynb) - nouveau \* \[GTE ModernBert\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/ModernBert.ipynb) - nouveau ### Grands LLM : \*\*Notebooks pour grands modèles :\*\* Ils dépassent le palier gratuit de 15 Go de VRAM de Colab. Avec les nouveaux GPU 80 Go de Colab, vous pouvez fine-tuner des modèles de 120 milliards de paramètres. {% hint style="info" %} Un abonnement ou des crédits Colab sont requis. Nous \*\*ne le faisons pas\*\* ne gagnons rien avec ces notebooks. {% endhint %} \* \[\*\*Gemma-4-31B\*\*\](https://www.kaggle.com/code/danielhanchen/gemma4-31b-unsloth) - nouveau et \*\*GRATUIT\*\* \* \[\*\*DiffusionGemma\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/DiffusionGemma\_\\(26B-A4B\\)-Sudoku.ipynb) - nouveau \* \[Gemma-4-26B-A4B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(26B\_A4B\\)-Vision.ipynb) - nouveau \* \[Gemma-4-31B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma4\_\\(31B\\)-Vision.ipynb) - nouveau \* \[Qwen3.5-35B-A3B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_MoE.ipynb) \* \[Qwen3.5‑27B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen\_3\_5\_27B\_A100\\(80GB\\).ipynb) \* \[GLM-4.7-Flash\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/GLM\_Flash\_A100\\(80GB\\).ipynb) \* \[gpt-oss-20b (contexte 500K)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_500K\_Context\_Fine\_tuning.ipynb) \* \[Qwen3-30B-A3B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_MoE.ipynb) \* \[Notebook LoRA Nemotron-3-Nano-30B-A3B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Nemotron-3-Nano-30B-A3B\_A100.ipynb) \* \[Notebook NeMo Gym Sudoku GRPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/NeMo-Gym-Sudoku.ipynb) \* \[Notebook NeMo Gym Multi Environment GRPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/NeMo-Gym-Multi-Environment.ipynb) \* \[gpt-oss-120b\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(120B\\)\_A100-Fine-tuning.ipynb) \* \[Qwen3 (32B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(32B\\)\_A100-Reasoning-Conversational.ipynb) \* \[Llama 3.3 (70B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.3\_\\(70B\\)\_A100-Conversational.ipynb) \* \[Gemma 3 (27B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(27B\\)\_A100-Conversational.ipynb) \* \[Baidu ERNIE 4.5 VL (28B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/ERNIE\_4\_5\_VL\_28B\_A3B\_PT\_Vision.ipynb) - nouveau ### Autres notebooks importants : \* \[\*\*Agent de support client\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Granite4.0.ipynb) \* \[Mistral Ministral 3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) - nouveau (résolution de sudoku) \* \[Déployer sur LM Studio \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\)-LMStudio.ipynb)- nouveau \* \[Entraînement sensible à la quantification\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)\_Instruct-QAT.ipynb) (QAT) - nouveau \* \[Déploiement sur téléphone \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(0\_6B\\)-Phone\_Deployment.ipynb)- nouveau \* \[Raisonner avant \*\*Appel d'outils\*\* notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\).ipynb) - nouveau \* \[notebook Mobile Actions\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\)-Mobile-Actions.ipynb) - nouveau \* \[\*\*Création automatique des kernels\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) avec RL \* \[\*\*ModernBERT-large\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert\_classification.ipynb) \*\*- nouveau\*\* 19 août \* \[\*\*Génération de données synthétiques Llama 3.2 (3B)\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Meta\_Synthetic\_Data\_Llama3\_2\_\\(3B\\).ipynb) \* \[gpt-oss-20b (contexte 500K)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_500K\_Context\_Fine\_tuning.ipynb) - nouveau (A100) \* \[\*\*Appel d'outils\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5\_Coder\_\\(1.5B\\)-Tool\_Calling.ipynb) \* \[Mistral v0.3 Instruct (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \[Ollama\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-Ollama.ipynb) \* \[ORPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-ORPO.ipynb) \* \[Pré-entraînement continu\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-CPT.ipynb) \* \[DPO Zephyr\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Zephyr\_\\(7B\\)-DPO.ipynb) \* \[\*\*\*Inférence uniquement\*\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1\_\\(8B\\)-Inference.ipynb) \* \[Llama 3 (8B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-Alpaca.ipynb) ### Notebooks pour cas d'utilisation spécifiques : \* \[Déploiement sur téléphone \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(0\_6B\\)-Phone\_Deployment.ipynb)- nouveau \* \[Déployer sur LM Studio \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\)-LMStudio.ipynb)- nouveau \* \[Raisonner avant \*\*Appel d'outils\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\).ipynb) - nouveau \* \[Mobile Actions\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\)-Mobile-Actions.ipynb) - nouveau \* \[\*\*Agent de support client\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Granite4.0.ipynb) \* \[Entraînement sensible à la quantification\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)\_Instruct-QAT.ipynb) (QAT) - nouveau \* \[\*\*Création automatique des kernels\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) avec RL \*\*- nouveau\*\* \* \[DPO Zephyr\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Zephyr\_\\(7B\\)-DPO.ipynb) \* \[BERT - classification de texte\](https://colab.research.google.com/github/timothelaborie/text\_classification\_scripts/blob/main/unsloth\_classification.ipynb) - (AutoModelForSequenceClassification) \* \[Ollama\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-Ollama.ipynb) \* \[\*\*Appel d'outils\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5\_Coder\_\\(1.5B\\)-Tool\_Calling.ipynb) \* \[Pré-entraînement continu (CPT)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-CPT.ipynb) \* \[Jeux de données multiples\](https://colab.research.google.com/drive/1njCCbE1YVal9xC83hjdo2hiGItpY\_D6t?usp=sharing) par Flail \* \[KTO\](https://colab.research.google.com/drive/1MRgGtLWuZX4ypSfGguFgC-IblTvO2ivM?usp=sharing) par Jeffrey \* \[Interface de chat d'inférence\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Unsloth\_Studio.ipynb) \* \[Conversationnel\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* \[ChatML\](https://colab.research.google.com/drive/15F1xyn8497\_dUbxZP4zWmPZ3PJx1Oymv?usp=sharing) \* \[Complétion de texte\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_\\(7B\\)-Text\_Completion.ipynb) ### Le reste des notebooks : \* \[Qwen2.5 (3B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5\_\\(3B\\)-GRPO.ipynb) \* \[Gemma 2 (9B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma2\_\\(9B\\)-Alpaca.ipynb) \* \[Mistral NeMo (12B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_Nemo\_\\(12B\\)-Alpaca.ipynb) \* \[Phi-3.5 (mini)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_3.5\_Mini-Conversational.ipynb) \* \[Phi-3 (medium)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_3\_Medium-Conversational.ipynb) \* \[Gemma 2 (2B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma2\_\\(2B\\)-Alpaca.ipynb) \* \[Qwen 2.5 Coder (14B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2.5\_Coder\_\\(14B\\)-Conversational.ipynb) \* \[Mistral Small (22B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_Small\_\\(22B\\)-Alpaca.ipynb) \* \[TinyLlama\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/TinyLlama\_\\(1.1B\\)-Alpaca.ipynb) \* \[CodeGemma (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/CodeGemma\_\\(7B\\)-Conversational.ipynb) \* \[Mistral v0.3 (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Alpaca.ipynb) \* \[Qwen2 (7B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2\_\\(7B\\)-Alpaca.ipynb) ## Notebooks Kaggle #### Notebooks standard : \* \[\*\*Gemma-4-31B\*\* (Kaggle)\](https://www.kaggle.com/code/danielhanchen/gemma4-31b-unsloth) - nouveau et \*\*GRATUIT\*\* \* \[\*\*gpt-oss (20B)\*\*\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-gpt-oss-\\(20B\\)-Fine-tuning.ipynb\\&accelerator=nvidiaTeslaT4) \* \[Gemma 3n (E4B)\](https://www.kaggle.com/code/danielhanchen/gemma-3n-4b-multimodal-finetuning-inference) \* \[Qwen3 (14B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Qwen3\_\\(14B\\).ipynb) \* \[Magistral-2509 (24B)\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Magistral\_\\(24B\\)-Reasoning-Conversational.ipynb\\&accelerator=nvidiaTeslaT4) \* \[Gemma 3 (4B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Gemma3\_\\(4B\\).ipynb) \* \[Phi-4 (14B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Phi\_4-Conversational.ipynb) \* \[Llama 3.1 (8B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3.1\_\\(8B\\)-Alpaca.ipynb) \* \[Llama 3.2 (1B + 3B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* \[Qwen 2.5 (7B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Qwen2.5\_\\(7B\\)-Alpaca.ipynb) #### Notebooks GRPO (raisonnement) : \* \[\*\*Qwen2.5-VL\*\*\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen2\_5\_7B\_VL\_GRPO.ipynb\\&accelerator=nvidiaTeslaT4) - GSPO Vision - nouveau \* \[Qwen3 (4B)\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen3\_\\(4B\\)-GRPO.ipynb\\&accelerator=nvidiaTeslaT4) \* \[Gemma 3 (1B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Gemma3\_\\(1B\\)-GRPO.ipynb) \* \[Llama 3.1 (8B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3.1\_\\(8B\\)-GRPO.ipynb) \* \[Phi-4 (14B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Phi\_4\_\\(14B\\)-GRPO.ipynb) \* \[Qwen 2.5 (3B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Qwen2.5\_\\(3B\\)-GRPO.ipynb) #### Notebooks de synthèse vocale (TTS) : \* \[Sesame-CSM (1B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Sesame\_CSM\_\\(1B\\)-TTS.ipynb) \* \[Orpheus-TTS (3B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Orpheus\_\\(3B\\)-TTS.ipynb) \* \[Whisper Large V3\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Whisper.ipynb) – reconnaissance vocale \* \[Llasa-TTS (1B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llasa\_TTS\_\\(1B\\).ipynb) \* \[Spark-TTS (0.5B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Spark\_TTS\_\\(0\_5B\\).ipynb) \* \[Oute-TTS (1B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Oute\_TTS\_\\(1B\\).ipynb) #### Notebooks Vision (multimodaux) : \* \[Llama 3.2 Vision (11B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3.2\_\\(11B\\)-Vision.ipynb) \* \[Qwen 2.5-VL (7B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Qwen2.5\_VL\_\\(7B\\)-Vision.ipynb) \* \[Pixtral (12B) 2409\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Pixtral\_\\(12B\\)-Vision.ipynb) #### Notebooks pour cas d'utilisation spécifiques : \* \[Appel d'outils\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen2.5\_Coder\_\\(1.5B\\)-Tool\_Calling.ipynb\\&accelerator=nvidiaTeslaT4) \* \[ORPO\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3\_\\(8B\\)-ORPO.ipynb) \* \[Pré-entraînement continu\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Mistral\_v0.3\_\\(7B\\)-CPT.ipynb) \* \[DPO Zephyr\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Zephyr\_\\(7B\\)-DPO.ipynb) \* \[Inférence uniquement\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3.1\_\\(8B\\)-Inference.ipynb) \* \[Ollama\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Llama3\_\\(8B\\)-Ollama.ipynb) \* \[Complétion de texte\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Mistral\_\\(7B\\)-Text\_Completion.ipynb) \* \[CodeForces-cot (raisonnement)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-CodeForces-cot-Finetune\_for\_Reasoning\_on\_CodeForces.ipynb) \* \[Unsloth Studio (interface de chat)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Unsloth\_Studio.ipynb) #### Le reste des notebooks : \* \[Gemma 2 (9B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Gemma2\_\\(9B\\)-Alpaca.ipynb) \* \[Gemma 2 (2B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Gemma2\_\\(2B\\)-Alpaca.ipynb) \* \[CodeGemma (7B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-CodeGemma\_\\(7B\\)-Conversational.ipynb) \* \[Mistral NeMo (12B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Mistral\_Nemo\_\\(12B\\)-Alpaca.ipynb) \* \[Mistral Small (22B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-Mistral\_Small\_\\(22B\\)-Alpaca.ipynb) \* \[TinyLlama (1.1B)\](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/Kaggle-TinyLlama\_\\(1.1B\\)-Alpaca.ipynb) Pour voir la liste complète de tous nos notebooks Kaggle, \[cliquez ici\](https://github.com/unslothai/notebooks#-kaggle-notebooks). {% hint style="info" %} N'hésitez pas à contribuer aux notebooks en visitant notre \[dépôt\](https://github.com/unslothai/notebooks)! {% endhint %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/unsloth-notebooks.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md). # Connect llama.cpp to Unsloth: Run GGUFs with llama-server Llama.cpp is an open-source inference engine for running GGUF models efficiently on local hardware, and \[Unsloth\](https://github.com/unslothai/unsloth) makes it easy to run those models directly into a open-source UI chat interface. By starting a local \`llama-server\`, you can serve a GGUF model from your machine or Hugging Face, connect it to Unsloth, and use it like any other external chat model. This guide walks through installing llama.cpp, launching \`llama-server\`, connecting it to Unsloth, enabling your model, and configuring prompt caching, context length, API keys, FA, and chat templates. ![](https://unsloth.ai/files/twmd9xWJs6E6cnUy3IEQ) \## Setup {% stepper %} {% step %} ### Install llama.cpp Install llama.cpp first so you can run the \`llama-server\` command. Use one of the official install options: \* Download a prebuilt \[llama.cpp binary\](https://github.com/ggml-org/llama.cpp/releases) \* Build llama.cpp from \[source\](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md) After installing, check that llama-server works in your terminal: \`llama-server --help\` {% endstep %} {% step %} ### Choose a GGUF model llama-server can load a local .gguf file or download a GGUF model from Hugging Face. To serve a Hugging Face GGUF repo directly, use the repo and quant name: \`llama-server -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4\_K\_XL\` If you wish to load a local model, you can also follow the steps below. Start \`llama-server\` with the model you want to serve: \`\`\`bash llama-server \\ --model /path/to/model.gguf \\ --host 0.0.0.0 \\ --port 8080 \`\`\` This exposes an API endpoint at: \`http://localhost:8080/v1\` To require an API key, add: \`\`\`bash --api-key 1234-myapi-key \`\`\` {% endstep %} {% step %} ### Connect Llama.cpp to Unsloth Open \*\*Settings → Connections\*\*, then click \*\*Add Connection\*\*. Select \*\*llama.cpp\*\*, then enter your server details: ![](https://unsloth.ai/files/hbAfB7WgobMottTyBHS0) if you did not start llama-server with \`--api-key\`, leave the API key field empty. Enter the base URL of your server, e.g. \`http://localhost:8080/v1\`\\ Click \*\*Load Models\*\* to fetch available model IDs, or enter model IDs manually if your server does not expose \`/models\`. ![](https://unsloth.ai/files/mNBYpzZH2VPxl4QRcXrg) Then, after you click \*\*Add Connection,\*\* The models you enabled will now appear under \*\*Connected\*\* in the \*\*Select Model\*\* dropdown. {% endstep %} {% step %} ### Ready to Chat After saving the connection, your llama.cpp model will appear under \*\*Connection\*\* in the model dropdown. Select it to start chatting through you \*\*llama-server\*\*. ![](https://unsloth.ai/files/twmd9xWJs6E6cnUy3IEQ) {% endstep %} {% endstepper %} ### Prompt Caching Prompt caching reduces latency and cost when requests reuse the same long prefix. Use the \*\*Prompt caching\*\* setting in the Unsloth side panel to control caching behaviour for supported connections. ![](https://unsloth.ai/files/rd7uqDkUz6YnddW01aRl) With llama.cpp, prompt caching is enabled by default and can be disabled when starting\\ \`llama-server\` with: \`\`\`bash --no-cache-prompt \`\`\` ### \*\*Common llama-server arguments\*\* The example above only uses the required connection settings. You can add more llama-server arguments depending on your model and hardware. Common options include: \`\`\`bash --ctx-size 8192 \\ # Set the context length --parallel 2 \\ # Set the number of parallel slots --flash-attn on \\ # Enable Flash Attention when supported --jinja \\ # Use the model chat template --api-key 1234-key \\ # Require an API key --no-cache-prompt # Disable prompt caching \`\`\` For the full list of server arguments, see the official \[llama.cpp server README\](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/phi-4-reasoning-how-to-run-and-fine-tune.md). # Phi-4 Reasoning : comment l'exécuter et le fine-tuner Les nouveaux modèles de raisonnement Phi-4 de Microsoft sont désormais pris en charge dans Unsloth. La variante « plus » affiche des performances comparables à celles de o1-mini, o3-mini et Sonnet 3.7 d’OpenAI. Les modèles de raisonnement « plus » et standard comptent 14B paramètres, tandis que le « mini » en a 4B.\\ \\ Tous les téléchargements de raisonnement Phi-4 utilisent notre \[Unsloth Dynamic 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) méthodologie. #### \*\*Raisonnement Phi-4 - téléchargements Unsloth Dynamic 2.0 :\*\* | Dynamic 2.0 GGUF (à exécuter) | Safetensor 4 bits dynamique (pour finetuner/déployer) | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | * [Raisonnement-plus](https://huggingface.co/unsloth/Phi-4-reasoning-plus-GGUF/) (14B) * [Raisonnement](https://huggingface.co/unsloth/Phi-4-reasoning-GGUF) (14B) * [Mini-raisonnement](https://huggingface.co/unsloth/Phi-4-mini-reasoning-GGUF/) (4B) | * [Raisonnement-plus](https://huggingface.co/unsloth/Phi-4-reasoning-plus-unsloth-bnb-4bit) * [Raisonnement](https://huggingface.co/unsloth/phi-4-reasoning-unsloth-bnb-4bit) * [Mini-raisonnement](https://huggingface.co/unsloth/Phi-4-mini-reasoning-unsloth-bnb-4bit) | ## 🖥️ \*\*Exécution du raisonnement Phi-4\*\* ### :gear: Paramètres officiels recommandés Selon Microsoft, voici les paramètres recommandés pour l’inférence : \* \*\*Température = 0,8\*\* \* Top\\\_P = 0.95 ### \*\*Modèles de chat Phi-4 reasoning\*\* Veuillez vous assurer d’utiliser le bon modèle de chat, car la variante « mini » en a un différent. #### \*\*Phi-4-mini :\*\* {% code overflow="wrap" %} \`\`\` <|system|>Votre nom est Phi, un expert en mathématiques IA développé par Microsoft.<|end|><|user|>Comment résoudre 3\*x^2+4\*x+5=1 ?<|end|><|assistant|> \`\`\` {% endcode %} #### \*\*Phi-4-reasoning et Phi-4-reasoning-plus :\*\* Ce format est utilisé pour la conversation générale et les instructions : {% code overflow="wrap" %} \`\`\` <|im\_start|>system<|im\_sep|>Vous êtes Phi, un modèle de langage entraîné par Microsoft pour aider les utilisateurs. Votre rôle d’assistant consiste à explorer minutieusement les questions au moyen d’un processus de réflexion systématique avant de fournir des solutions finales précises et exactes. Cela nécessite de s’engager dans un cycle complet d’analyse, de synthèse, d’exploration, de réévaluation, de réflexion, de retour en arrière et d’itération afin de développer un processus de réflexion bien considéré. Veuillez structurer votre réponse en deux sections principales : Thought et Solution, en utilisant le format spécifié : {Section Thought} {Section Solution}. Dans la section Thought, détaillez votre processus de raisonnement par étapes. Chaque étape doit inclure des considérations détaillées telles que l’analyse des questions, la synthèse des résultats pertinents, le brainstorming de nouvelles idées, la vérification de l’exactitude des étapes actuelles, l’affinement des éventuelles erreurs et la révision des étapes précédentes. Dans la section Solution, en vous basant sur les différentes tentatives, explorations et réflexions de la section Thought, présentez systématiquement la solution finale que vous jugez correcte. La section Solution doit être logique, précise et concise, et détailler les étapes nécessaires pour parvenir à la conclusion. Maintenant, essayez de résoudre la question suivante en suivant les consignes ci-dessus :<|im\_end|><|im\_start|>user<|im\_sep|>Que vaut 1+1 ?<|im\_end|><|im\_start|>assistant<|im\_sep|> \`\`\` {% endcode %} {% hint style="info" %} Oui, le format du modèle de chat/prompt est aussi long que cela ! {% endhint %} ### 🦙 Ollama : tutoriel pour exécuter le raisonnement Phi-4 1. Installez \`ollama\` si vous ne l’avez pas déjà fait ! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Exécutez le modèle ! Notez que vous pouvez appeler \`ollama serve\`dans un autre terminal si cela échoue. Nous incluons toutes nos corrections et les paramètres suggérés (température, etc.) dans \`params\` dans notre dépôt sur Hugging Face. \`\`\`bash ollama run hf.co/unsloth/Phi-4-mini-reasoning-GGUF:Q4\_K\_XL \`\`\` ### 📖 Llama.cpp : tutoriel pour exécuter le raisonnement Phi-4 {% hint style="warning" %} Vous devez utiliser \`--jinja\` dans llama.cpp pour activer le raisonnement pour les modèles, sauf pour la variante « mini ». Sinon, aucun token ne sera fourni. {% endhint %} 1. Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir Q4\\\_K\\\_M, ou d’autres versions quantifiées. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Phi-4-mini-reasoning-GGUF", local\_dir = "unsloth/Phi-4-mini-reasoning-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` 3. Exécutez le modèle en mode conversationnel dans llama.cpp. Vous devez utiliser \`--jinja\` dans llama.cpp pour activer le raisonnement pour les modèles. Cela n’est toutefois pas nécessaire si vous utilisez la variante « mini ». \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Phi-4-mini-reasoning-GGUF/Phi-4-mini-reasoning-UD-Q4\_K\_XL.gguf \\ --threads -1 \\\\ --n-gpu-layers 99 \\ --prio 3 \\ --temp 0.8 \\ --top-p 0.95 \\ --jinja \\ --min-p 0.00 \\ --ctx-size 32768 \\\\ --seed 3407 \`\`\` ## 🦥 Fine-tuning de Phi-4 avec Unsloth \[Fine-tuning de Phi-4\](https://unsloth.ai/blog/phi4) pour les modèles sont également désormais pris en charge dans Unsloth. Pour fine-tuner gratuitement sur Google Colab, il suffit de modifier le \`model\_name\` de 'unsloth/Phi-4' à 'unsloth/Phi-4-mini-reasoning', etc. \* \[carnet de fine-tuning de Phi-4 (14B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/phi-4-reasoning-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md). # Gemma 3 - How to Run Guide Google releases Gemma 3 with a new 270M model and the previous 1B, 4B, 12B, and 27B sizes. The 270M and 1B are text-only, while larger models handle both text and vision. We provide GGUFs, and a guide of how to run it effectively, and how to finetune & do \[RL\](/docs/get-started/reinforcement-learning-rl-guide.md) with Gemma 3! {% hint style="success" %} \*\*NEW Aug 14, 2025 Update:\*\* Try our fine-tuning \[Gemma 3 (270M) notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(270M\\).ipynb) and \[GGUFs to run\](https://huggingface.co/collections/unsloth/gemma-3-67d12b7e8816ec6efa7e4e5b). Also see our \[Gemma 3n Guide\](/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune/gemma-3n-how-to-run-and-fine-tune.md). {% endhint %} [Running Tutorial](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md#gmail-running-gemma-3-on-your-phone) [Fine-tuning Tutorial](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md#fine-tuning-gemma-3-in-unsloth) \*\*Unsloth is the only framework which works in float16 machines for Gemma 3 inference and training.\*\* This means Colab Notebooks with free Tesla T4 GPUs also work! \* Fine-tune Gemma 3 (4B) with vision support using our \[free Colab notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) {% hint style="info" %} According to the Gemma team, the optimal config for inference is\\ \`temperature = 1.0, top\_k = 64, top\_p = 0.95, min\_p = 0.0\` {% endhint %} \*\*Unsloth Gemma 3 uploads with optimal configs:\*\* | GGUF | Unsloth Dynamic 4-bit Instruct | 16-bit Instruct | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it-GGUF) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b-it-GGUF) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b-it-GGUF) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b-it-GGUF) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b-it-GGUF) | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it-unsloth-bnb-4bit) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b-it-bnb-4bit) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b-it-bnb-4bit) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b-it-unsloth-bnb-4bit) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b-it-bnb-4bit) | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b) | ## :gear: Recommended Inference Settings According to the Gemma team, the official recommended settings for inference is: \* Temperature of 1.0 \* Top\\\_K of 64 \* Min\\\_P of 0.00 (optional, but 0.01 works well, llama.cpp default is 0.1) \* Top\\\_P of 0.95 \* Repetition Penalty of 1.0. (1.0 means disabled in llama.cpp and transformers) \* Chat template: user\nHello!\nmodel\nHey there!\nuser\nWhat is 1+1?\nmodel\n \* Chat template with \`\\n\`newlines rendered (except for the last) {% code overflow="wrap" %} \`\`\` user Hello! model Hey there! user What is 1+1? model\\n \`\`\` {% endcode %} {% hint style="danger" %} llama.cpp an other inference engines auto add a \\ - DO NOT add TWO \\ tokens! You should ignore the \\ when prompting the model! {% endhint %} ### ✨Running Gemma 3 on your phone [](https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md#gmail-running-gemma-3-on-your-phone) To run the models on your phone, we recommend using any mobile app that can run GGUFs locally on edge devices like phones. After fine-tuning you can export it to GGUF then run it locally on your phone. Ensure your phone has enough RAM/power to process the models as it can overheat so we recommend using Gemma 3 270M or the Gemma 3n models for this use-case. You can try the \[open-source project AnythingLLM's\](https://github.com/Mintplex-Labs/anything-llm) mobile app which you can download on \[Android here\](https://play.google.com/store/apps/details?id=com.anythingllm) or \[ChatterUI\](https://github.com/Vali-98/ChatterUI), which are great apps for running GGUFs on your phone. {% hint style="success" %} Remember, you can change the model name 'gemma-3-27b-it-GGUF' to any Gemma model like 'gemma-3-270m-it-GGUF:Q8\\\_K\\\_XL' for all the tutorials. {% endhint %} ## :llama: Tutorial: How to Run Gemma 3 in Ollama 1. Install \`ollama\` if you haven't already! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Run the model! Note you can call \`ollama serve\`in another terminal if it fails! We include all our fixes and suggested parameters (temperature etc) in \`params\` in our Hugging Face upload! You can change the model name 'gemma-3-27b-it-GGUF' to any Gemma model like 'gemma-3-270m-it-GGUF:Q8\\\_K\\\_XL'. \`\`\`bash ollama run hf.co/unsloth/gemma-3-27b-it-GGUF:Q4\_K\_XL \`\`\` ## 📖 Tutorial: How to Run Gemma 3 27B in llama.cpp 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=ON -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. If you want to use \`llama.cpp\` directly to load models, you can do the below: (:Q4\\\_K\\\_XL) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` \`\`\`bash ./llama.cpp/llama-mtmd-cli \\ -hf unsloth/gemma-3-4b-it-GGUF:Q4\_K\_XL \`\`\` 3. \*\*OR\*\* download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose Q4\\\_K\\\_M, or other quantized versions (like BF16 full precision). More versions at: \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/gemma-3-27b-it-GGUF", local\_dir = "unsloth/gemma-3-27b-it-GGUF", allow\_patterns = \["\*Q4\_K\_XL\*", "mmproj-BF16.gguf"\], # For Q4\_K\_M ) \`\`\` 4. Run Unsloth's Flappy Bird test 5. Edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length (Gemma 3 supports 128K context length!), \`--n-gpu-layers 99\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. 6. For conversation mode: \`\`\`bash ./llama.cpp/llama-mtmd-cli \\ --model unsloth/gemma-3-27b-it-GGUF/gemma-3-27b-it-Q4\_K\_XL.gguf \\ --mmproj unsloth/gemma-3-27b-it-GGUF/mmproj-BF16.gguf \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 1.0 \\ --repeat-penalty 1.0 \\ --min-p 0.01 \\ --top-k 64 \\ --top-p 0.95 \`\`\` 7. For non conversation mode to test Flappy Bird: \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/gemma-3-27b-it-GGUF/gemma-3-27b-it-Q4\_K\_XL.gguf \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 1.0 \\ --repeat-penalty 1.0 \\ --min-p 0.01 \\ --top-k 64 \\ --top-p 0.95 \\ -no-cnv \\ --prompt "user\\nCreate a Flappy Bird game in Python. You must include these things:\\n1. You must use pygame.\\n2. The background color should be randomly chosen and is a light shade. Start with a light blue color.\\n3. Pressing SPACE multiple times will accelerate the bird.\\n4. The bird's shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color.\\n5. Place on the bottom some land colored as dark brown or yellow chosen randomly.\\n6. Make a score shown on the top right side. Increment if you pass pipes and don't hit them.\\n7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade.\\n8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again.\\nThe final game should be inside a markdown section in Python. Check your code for errors and fix them before the final markdown section.\\nmodel\\n" \`\`\` The full input from our 1.58bit blog is: {% hint style="danger" %} Remember to remove \\ since Gemma 3 auto adds a \\! {% endhint %} {% code overflow="wrap" %} \`\`\` user Create a Flappy Bird game in Python. You must include these things: 1. You must use pygame. 2. The background color should be randomly chosen and is a light shade. Start with a light blue color. 3. Pressing SPACE multiple times will accelerate the bird. 4. The bird's shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color. 5. Place on the bottom some land colored as dark brown or yellow chosen randomly. 6. Make a score shown on the top right side. Increment if you pass pipes and don't hit them. 7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade. 8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again. The final game should be inside a markdown section in Python. Check your code for error \`\`\` {% endcode %} ## :sloth: Fine-tuning Gemma 3 in Unsloth \*\*Unsloth is the only framework which works in float16 machines for Gemma 3 inference and training.\*\* This means Colab Notebooks with free Tesla T4 GPUs also work! \* Try our new \[Gemma 3 (270M) notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(270M\\).ipynb) which makes the 270M parameter model very smart at playing chess and can predict the next chess move. \* Fine-tune Gemma 3 (4B) using our notebooks for: \[\*\*Text\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) or \[\*\*Vision\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) \* Or fine-tune \[Gemma 3n (E4B)\](/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune/gemma-3n-how-to-run-and-fine-tune.md) with \[Text\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Conversational.ipynb) • \[Vision\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Vision.ipynb) • \[Audio\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Audio.ipynb) {% hint style="warning" %} When trying full fine-tune (FFT) Gemma 3, all layers default to float32 on float16 devices. Unsloth expects float16 and upcasts dynamically. To fix, run \`model.to(torch.float16)\` after loading, or use a GPU with bfloat16 support. {% endhint %} ### Unsloth Fine-tuning Fixes Our solution in Unsloth is 3 fold: 1. Keep all intermediate activations in bfloat16 format - can be float32, but this uses 2x more VRAM or RAM (via Unsloth's async gradient checkpointing) 2. Do all matrix multiplies in float16 with tensor cores, but manually upcasting / downcasting without the help of Pytorch's mixed precision autocast. 3. Upcast all other options that don't need matrix multiplies (layernorms) to float32. ## 🤔 Gemma 3 Fixes Analysis ![](https://unsloth.ai/files/W5Sspypmy4p2Pii0h8mp) Gemma 3 1B to 27B exceed float16's maximum of 65504 First, before we finetune or run Gemma 3, we found that when using float16 mixed precision, gradients and \*\*activations become infinity\*\* unfortunately. This happens in T4 GPUs, RTX 20x series and V100 GPUs where they only have float16 tensor cores. For newer GPUs like RTX 30x or higher, A100s, H100s etc, these GPUs have bfloat16 tensor cores, so this problem does not happen! \*\*But why?\*\* ![](https://unsloth.ai/files/zZuoVdQ70rKsVRZiMrzz) Wikipedia [https://en.wikipedia.org/wiki/Bfloat16\_floating-point\_format](https://en.wikipedia.org/wiki/Bfloat16_floating-point_format) Float16 can only represent numbers up to \*\*65504\*\*, whilst bfloat16 can represent huge numbers up to \*\*10^38\*\*! But notice both number formats use only 16bits! This is because float16 allocates more bits so it can represent smaller decimals better, whilst bfloat16 cannot represent fractions well. But why float16? Let's just use float32! But unfortunately float32 in GPUs is very slow for matrix multiplications - sometimes 4 to 10x slower! So we cannot do this. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md). # DeepSeek-OCR : comment l'exécuter et le fine-tuner \*\*DeepSeek-OCR\*\* est un modèle de vision de 3 milliards de paramètres pour la reconnaissance optique de caractères (OCR) et la compréhension de documents. Il utilise \*compression optique contextuelle\* pour convertir des mises en page 2D en tokens visuels, permettant un traitement efficace des contextes longs. Capable de gérer les tableaux, les articles et l'écriture manuscrite, DeepSeek-OCR atteint 97 % de précision tout en utilisant 10× moins de tokens visuels que de tokens textuels - ce qui le rend 10× plus efficace que les LLM basés sur du texte. Vous pouvez affiner DeepSeek-OCR pour améliorer ses performances visuelles ou linguistiques. Dans notre Unsloth \[\*\*carnet d'affinage gratuit\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb), nous avons démontré une \[amélioration de 88,26 %\](#fine-tuning-deepseek-ocr) pour la compréhension du langage. [Exécution de DeepSeek-OCR](https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md#running-deepseek-ocr) [Affinage de DeepSeek-OCR](https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md#fine-tuning-deepseek-ocr) > \*\*Notre upload de modèle qui permet l'affinage + plus de support d'inférence :\*\* \[\*\*DeepSeek-OCR\*\*\](https://huggingface.co/unsloth/DeepSeek-OCR) ## 🖥️ \*\*Exécution de DeepSeek-OCR\*\* Pour exécuter le modèle dans \[vLLM\](#vllm-run-deepseek-ocr-tutorial) ou \[Unsloth\](#unsloth-run-deepseek-ocr-tutorial), voici les paramètres recommandés : ### :gear: Paramètres recommandés DeepSeek recommande ces paramètres : \* \*\*Température = 0.0\*\* \* \`max\_tokens = 8192\` \* \`ngram\_size = 30\` \* \`window\_size = 90\` ### 📖 vLLM : Tutoriel d'exécution de DeepSeek-OCR 1. Obtenez le dernier \`vLLM\` via : \`\`\`bash uv venv source .venv/bin/activate # Jusqu'à la version v0.11.1, vous devez installer vLLM depuis la build nightly uv pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightly \`\`\` 2. Ensuite, exécutez le code suivant : {% code overflow="wrap" %} \`\`\`python from vllm import LLM, SamplingParams from vllm.model\_executor.models.deepseek\_ocr import NGramPerReqLogitsProcessor from PIL import Image # Créer une instance du modèle llm = LLM( model="unsloth/DeepSeek-OCR", enable\_prefix\_caching=False, mm\_processor\_cache\_gb=0, logits\_processors=\[NGramPerReqLogitsProcessor\], ) # Préparer une entrée par lot avec votre fichier image image\_1 = Image.open("path/to/your/image\_1.png").convert("RGB") image\_2 = Image.open("path/to/your/image\_2.png").convert("RGB") prompt = "\\nOCR libre." model\_input = \[ { "prompt": prompt, "multi\_modal\_data": {"image": image\_1} }, { "prompt": prompt, "multi\_modal\_data": {"image": image\_2} } \] sampling\_param = SamplingParams( temperature=0.0, max\_tokens=8192, # arguments du processeur de logits ngram extra\_args=dict( ngram\_size=30, window\_size=90, whitelist\_token\_ids={128821, 128822}, # liste blanche : , ), skip\_special\_tokens=False, ) # Générer la sortie model\_outputs = llm.generate(model\_input, sampling\_param) # Imprimer la sortie for output in model\_outputs: print(output.outputs\[0\].text) \`\`\` {% endcode %} ### 🦥 Unsloth : Tutoriel d'exécution de DeepSeek-OCR 1. Obtenez le dernier \`unsloth\` via \`pip install --upgrade unsloth\` . Si vous avez déjà Unsloth, mettez-le à jour via \`pip install --upgrade --force-reinstall --no-deps --no-cache-dir unsloth unsloth\_zoo\` 2. Ensuite, utilisez le code ci-dessous pour exécuter DeepSeek-OCR : {% code overflow="wrap" %} \`\`\`python from unsloth import FastVisionModel import torch from transformers import AutoModel import os os.environ\["UNSLOTH\_WARN\_UNINITIALIZED"\] = '0' from huggingface\_hub import snapshot\_download snapshot\_download("unsloth/DeepSeek-OCR", local\_dir = "deepseek\_ocr") model, tokenizer = FastVisionModel.from\_pretrained( "./deepseek\_ocr", load\_in\_4bit = False, # Use 4bit to reduce memory use. False for 16bit LoRA. auto\_model = AutoModel, trust\_remote\_code = True, unsloth\_force\_compile = True, use\_gradient\_checkpointing = "unsloth", # True or "unsloth" for long context ) prompt = "\\nFree OCR. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 640, crop\_mode=True, save\_results = True, test\_compress = False) \`\`\` {% endcode %} ## 🦥 \*\*Affinage de DeepSeek-OCR\*\* Unsloth prend en charge l'affinage de DeepSeek-OCR. Étant donné que le modèle par défaut n'est pas exécutable sur la dernière \`transformers\` version, nous avons ajouté les modifications de l'équipe \[Stranger Vision HF\](https://huggingface.co/strangervisionhf) pour ensuite permettre l'inférence. Comme d'habitude, Unsloth entraîne DeepSeek-OCR 1,4× plus vite avec 40 % de VRAM en moins et des longueurs de contexte 5× plus grandes - sans dégradation de la précision.\\ \\ Nous avons créé deux notebooks Colab DeepSeek-OCR gratuits (avec et sans évaluation) : \* DeepSeek-OCR : \[Carnet uniquement pour l'affinage\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb) \* DeepSeek-OCR : \[Notebook d'affinage + évaluation\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\)-Eval.ipynb) (A100) L'affinage de DeepSeek-OCR sur un échantillon de 200K en persan a entraîné des gains substantiels dans la détection et la compréhension du texte persan. Nous avons évalué le modèle de base par rapport à notre version affinée sur 200 échantillons de transcriptions persanes, observant une \*\*amélioration absolue de 88,26 %\*\* dans le taux d'erreur de caractères (CER). Après seulement 60 étapes d'entraînement (taille de lot = 8), le CER moyen a diminué de \*\*149.07%\*\* pour atteindre une moyenne de \*\*60.81%\*\*. Cela signifie que le modèle affiné est \*\*57%\*\* plus précis pour comprendre le persan. Vous pouvez remplacer le jeu de données persan par le vôtre pour améliorer DeepSeek-OCR pour d'autres cas d'utilisation.\\ \\ Pour les résultats d'éval replica-table, utilisez notre notebook d'éval ci-dessus. Pour des résultats d'éval détaillés, voir ci-dessous : ### Résultats de l'évaluation du modèle affiné : {% columns fullWidth="true" %} {% column %} \*\*DeepSeek-OCR Baseline\*\* Performance moyenne du modèle de base : 149,07 % de CER pour cet ensemble d'évaluation ! \`\`\` ============================================================ Performance du modèle de base ============================================================ Nombre d'échantillons : 200 CER moyen : 149,07 % CER médian : 80,00 % Écart-type : 310,39 % CER min : 0,00 % CER max : 3500,00 % ============================================================ Meilleures prédictions (CER le plus bas) : Échantillon 5024 (CER : 0,00 %) Référence : چون هستی خیلی زیاد... Prédiction : چون هستی خیلی زیاد... Échantillon 3517 (CER : 0,00 %) Référence : تو ایران هیچوقت از اینها وجود نخواهد داشت... Prédiction : تو ایران هیچوقت از اینها وجود نخواهد داشت... Échantillon 9949 (CER : 0,00 %) Référence : کاش میدونستم هیچی بیخیال... Prédiction : کاش میدونستم هیچی بیخیال... Pires prédictions (CER le plus élevé) : Échantillon 11155 (CER : 3500,00 %) Référence : خسو... Prédiction : \\\[ \\text{CH}\_3\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}\_2\\text{CH}... Échantillon 13366 (CER : 1900,00 %) Référence : مشو... Prédiction : \\\[\\begin{align\*}\\underline{\\mathfrak{su}}\_0\\end{align\*}\\\]... Échantillon 10552 (CER : 1014,29 %) Référence : هیییییچ... Prédiction : e \`\`\` {% endcolumn %} {% column %} \*\*DeepSeek-OCR Affiné\*\* En 60 étapes, nous avons réduit le CER de 149,07 % à 60,43 % (amélioration du CER de 89 %)\ \ ============================================================\ Performance du modèle affiné\ ============================================================\ Nombre d'échantillons : 200\ CER moyen : 60,43 %\ CER médian : 50,00 %\ Écart-type : 80,63 %\ CER min : 0,00 %\ CER max : 916,67 %\ ============================================================\ \ Meilleures prédictions (CER le plus bas) :\ \ Échantillon 301 (CER : 0,00 %)\ Référence : باشه بابا تو لاکچری، تو خاص، تو خفن...\ Prédiction : باشه بابا تو لاکچری، تو خاص، تو خفن...\ \ Échantillon 2512 (CER : 0,00 %)\ Référence : از شخص حاج عبدالله زنجبیلی میگیرنش...\ Prédiction : از شخص حاج عبدالله زنجبیلی میگیرنش...\ \ Échantillon 2713 (CER : 0,00 %)\ Référence : نمی دونم والا تحمل نقد ندارن ظاهرا...\ Prédiction : نمی دونم والا تحمل نقد ندارن ظاهرا...\ \ Pires prédictions (CER le plus élevé) :\ \ Échantillon 14270 (CER : 916,67 %)\ Référence : ۴۳۵۹۴۷۴۷۳۸۹۰...\ Prédiction : پروپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپریپیپریپریپریپریپریپریپریپریپریپریپریپر...\ \ Échantillon 3919 (CER : 380,00 %)\ Référence : ۷۵۵۰۷۱۰۶۵۹...\ Prédiction : وادووووووووووووووووووووووووووووووووووو...\ \ Échantillon 3718 (CER : 333,33 %)\ Référence : ۳۲۶۷۲۲۶۵۵۸۴۶...\ Prédiction : پُپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُسوپُ...\ \ \ {% endcolumn %} {% endcolumns %} Un exemple tiré du jeu de données persan de 200K que nous avons utilisé (vous pouvez utiliser le vôtre), montrant l'image à gauche et le texte correspondant à droite.\ \ ![](https://unsloth.ai/files/a49132d1a1f677fd064e49b5d6bcc1f67cf0537b)\ \ \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/grok-2.md). # Grok 2 Vous pouvez maintenant exécuter \*\*Grok 2\*\* (alias Grok 2.5), le modèle à 270B de paramètres de xAI. La précision complète nécessite \*\*539 Go\*\*, tandis que la version dynamique 3 bits d’Unsloth réduit la taille à seulement \*\*118 Go\*\* (une réduction de 75 %). GGUF : \[Grok-2-GGUF\](https://huggingface.co/unsloth/grok-2-GGUF) Le \*\*Q3\\\_K\\\_XL 3 bits\*\* modèle fonctionne sur un seul \*\*Mac 128 Go\*\* ou \*\*24 Go de VRAM + 128 Go de RAM\*\*, atteignant \*\*plus de 5 jetons/s\*\* d’inférence. Merci à l’équipe de llama.cpp et à la communauté pour \[la prise en charge de Grok 2\](https://github.com/ggml-org/llama.cpp/pull/15539) et pour avoir rendu cela possible. Nous avons aussi été heureux d’avoir pu apporter un petit coup de main en chemin ! Tous les téléchargements utilisent Unsloth \[Dynamic 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) pour des performances de pointe en MMLU 5-shot et en divergence KL, ce qui signifie que vous pouvez exécuter des LLM Grok quantifiés avec une perte de précision minimale. [Tutoriel d’exécution dans llama.cpp](https://unsloth.ai/docs/fr/modeles/tutorials/grok-2.md#run-in-llama.cpp) ## :gear: Paramètres recommandés La quantification dynamique 3 bits utilise 118 Go (126 Gio) d’espace disque — cela fonctionne bien sur un Mac avec mémoire unifiée de 128 Go de RAM ou sur une carte 1x24 Go avec 128 Go de RAM. Il est recommandé d’avoir au moins 120 Go de RAM pour exécuter cette quantification 3 bits. {% hint style="warning" %} Vous devez utiliser \`--jinja\` pour Grok 2. Vous pourriez obtenir des résultats incorrects si vous n’utilisez pas \`--jinja\` {% endhint %} La quantification 8 bits fait environ 300 Go et tient sur un GPU 1x 80 Go (avec les couches MoE déchargées vers la RAM). Attendez-vous à environ 5 jetons/s avec cette configuration si vous disposez également de 200 Go de RAM supplémentaires. Pour apprendre à augmenter la vitesse de génération et à prendre en charge des contextes plus longs, \[lisez ici\](#improving-generation-speed). {% hint style="info" %} Bien que ce ne soit pas indispensable, pour de meilleures performances, veillez à ce que votre VRAM + RAM combinées soient égales à la taille de la quantification que vous téléchargez. Sinon, le déchargement vers le disque dur / SSD fonctionnera avec llama.cpp, mais l’inférence sera plus lente. {% endhint %} ### Paramètres d’échantillonnage \* Grok 2 a une longueur de contexte maximale de 128K, donc utilisez \`131,072\` de contexte ou moins. \* Utilisez \`--jinja\` pour les variantes de llama.cpp Il n’existe pas de paramètres d’échantillonnage officiels pour exécuter le modèle, vous pouvez donc utiliser les valeurs par défaut standard pour la plupart des modèles : \* Définissez le \*\*température = 1.0\*\* \* \*\*Min\\\_P = 0.01\*\* (facultatif, mais 0.01 fonctionne bien ; la valeur par défaut de llama.cpp est 0.1) ## Tutoriel d’exécution de Grok 2 : Actuellement, vous ne pouvez exécuter Grok 2 que dans llama.cpp. ### ✨ Exécuter dans llama.cpp {% stepper %} {% step %} Installez le \`llama.cpp\` PR spécifique pour Grok 2 sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp/pull/15539). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cd llama.cpp && git fetch origin pull/15539/head:MASTER && git checkout MASTER && cd .. cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli llama-server cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} Si vous souhaitez utiliser \`llama.cpp\` directement pour charger des modèles, vous pouvez faire ce qui suit : (:Q3\\\_K\\\_XL) est le type de quantification. Vous pouvez aussi télécharger via Hugging Face (point 3). C’est similaire à \`ollama run\` . Utilisez \`export LLAMA\_CACHE="folder"\` pour forcer \`llama.cpp\` à être enregistré à un emplacement spécifique. N'oubliez pas que le modèle a une longueur de contexte maximale de 128K uniquement. {% hint style="info" %} Veuillez essayer \`-ot ".ffn\_.\*\_exps.=CPU"\` pour décharger toutes les couches MoE vers le CPU ! Cela permet effectivement de faire tenir toutes les couches non MoE sur 1 GPU, améliorant ainsi les vitesses de génération. Vous pouvez personnaliser l'expression regex pour faire tenir davantage de couches si vous disposez de plus de capacité GPU. Si vous avez un peu plus de mémoire GPU, essayez \`-ot ".ffn\_(up|down)\_exps.=CPU"\` Cela décharge les couches MoE de projection montante et descendante. Essayez \`-ot ".ffn\_(up)\_exps.=CPU"\` si vous avez encore plus de mémoire GPU. Cela décharge uniquement les couches MoE de projection montante. Et enfin, déchargez toutes les couches via \`-ot ".ffn\_.\*\_exps.=CPU"\` Cela utilise le moins de VRAM. Vous pouvez aussi personnaliser la regex, par exemple \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` signifie décharger les couches MoE gate, up et down, mais uniquement à partir de la 6e couche. {% endhint %} \`\`\`bash export LLAMA\_CACHE="unsloth/grok-2-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/grok-2-GGUF:Q3\_K\_XL \\ --jinja \\ --n-gpu-layers 99 \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` {% endstep %} {% step %} Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir \`UD-Q3\_K\_XL\` (quantification dynamique 3 bits) ou d’autres versions quantifiées comme \`Q4\_K\_M\` . Nous \*\*recommandons d’utiliser notre quantification dynamique 2,7 bits\*\*\*\* \*\*\*\*\`UD-Q2\_K\_XL\`\*\*\*\* \*\*\*\*ou plus pour équilibrer la taille et la précision\*\*. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "0" # Peut parfois entraîner une limitation de débit, donc mettre à 0 pour désactiver from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/grok-2-GGUF", local\_dir = "unsloth/grok-2-GGUF", allow\_patterns = \["\*UD-Q3\_K\_XL\*"\], # 3 bits dynamique ) \`\`\` {% endstep %} {% step %} Vous pouvez modifier \`--threads 32\` pour le nombre de threads CPU, \`--ctx-size 16384\` pour la longueur du contexte, \`--n-gpu-layers 2\` pour le déchargement GPU, selon le nombre de couches. Essayez de l’ajuster si votre GPU manque de mémoire. Supprimez-le aussi si vous n'avez qu'une inférence CPU. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/grok-2-GGUF/UD-Q3\_K\_XL/grok-2-UD-Q3\_K\_XL-00001-of-00003.gguf \\ --jinja \\ --threads -1 \\\\ --n-gpu-layers 99 \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` {% endcode %} {% endstep %} {% endstepper %} ## Téléversements du modèle \*\*TOUS nos téléversements\*\* - y compris ceux qui ne sont pas basés sur imatrix ou dynamiques - utilisent notre jeu de données de calibration, spécialement optimisé pour les tâches conversationnelles, de codage et de langage. | Bits MoE | Type + lien | Taille sur disque | Détails | | -------- | ----------------------------------------------------------------------------------- | ----------------- | ------------- | | 1,66 bit | \[TQ1\\\_0\](https://huggingface.co/unsloth/grok-2-GGUF/blob/main/grok-2-UD-TQ1\_0.gguf) | \*\*81,8 Go\*\* | 1,92/1,56 bit | | 1,78 bit | \[IQ1\\\_S\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ1\_S) | \*\*88,9 Go\*\* | 2,06/1,56 bit | | 1,93 bit | \[IQ1\\\_M\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ1\_M) | \*\*94,5 Go\*\* | 2.5/2.06/1.56 | | 2,42 bit | \[IQ2\\\_XXS\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ2\_XXS) | \*\*99,3 Go\*\* | 2,5/2,06 bit | | 2,71 bit | \[Q2\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q2\_K\_XL) | \*\*112 Go\*\* | 3,5/2,5 bit | | 3,12 bit | \[IQ3\\\_XXS\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ3\_XXS) | \*\*117 Go\*\* | 3,5/2,06 bit | | 3,5 bit | \[Q3\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q3\_K\_XL) | \*\*126 Go\*\* | 4,5/3,5 bit | | 4,5 bit | \[Q4\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q4\_K\_XL) | \*\*155 Go\*\* | 5,5/4,5 bit | | 5,5 bit | \[Q5\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q5\_K\_XL) | \*\*191 Go\*\* | 6,5/5,5 bit | ## :snowboarder: Améliorer la vitesse de génération Si vous avez plus de VRAM, vous pouvez essayer de décharger davantage de couches MoE, ou de décharger des couches entières. Normalement, \`-ot ".ffn\_.\*\_exps.=CPU"\` décharge toutes les couches MoE vers le CPU ! Cela permet effectivement de faire tenir toutes les couches non MoE sur 1 GPU, améliorant ainsi les vitesses de génération. Vous pouvez personnaliser l'expression regex pour faire tenir davantage de couches si vous disposez de plus de capacité GPU. Si vous avez un peu plus de mémoire GPU, essayez \`-ot ".ffn\_(up|down)\_exps.=CPU"\` Cela décharge les couches MoE de projection montante et descendante. Essayez \`-ot ".ffn\_(up)\_exps.=CPU"\` si vous avez encore plus de mémoire GPU. Cela décharge uniquement les couches MoE de projection montante. Vous pouvez aussi personnaliser la regex, par exemple \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` signifie décharger les couches MoE gate, up et down, mais uniquement à partir de la 6e couche. Le \[la dernière version de llama.cpp\](https://github.com/ggml-org/llama.cpp/pull/14363) introduit également le mode à haut débit. Utilisez \`llama-parallel\`. En savoir plus \[ici\](https://github.com/ggml-org/llama.cpp/tree/master/examples/parallel). Vous pouvez aussi \*\*quantifier le cache KV en 4 bits\*\* par exemple pour réduire les déplacements VRAM / RAM, ce qui peut aussi accélérer le processus de génération. ## 📐Comment gérer un long contexte (128K complet) Pour faire tenir un contexte plus long, vous pouvez utiliser \*\*la quantification du cache KV\*\* pour quantifier les caches K et V en moins de bits. Cela peut aussi augmenter la vitesse de génération grâce à la réduction des transferts de données RAM / VRAM. Les options autorisées pour la quantification K (la valeur par défaut est \`f16\`) sont les suivantes. \`--cache-type-k f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` Vous devriez utiliser les \`\_1\` variantes pour une précision légèrement meilleure, bien que ce soit un peu plus lent. Par exemple \`q4\_1, q5\_1\` Vous pouvez aussi quantifier le cache V, mais vous devrez \*\*compiler llama.cpp avec la prise en charge de Flash Attention\*\* via \`-DGGML\_CUDA\_FA\_ALL\_QUANTS=ON\`, et utiliser \`--flash-attn\` pour l’activer. Ensuite, vous pouvez l'utiliser avec \`--cache-type-k\` : \`--cache-type-v f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/grok-2.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/ibm-granite-4.0.md). # IBM Granite 4.0 IBM lance les modèles Granite-4.0 avec 3 tailles, dont \*\*Nano\*\* (350M et 1B), \*\*Micro\*\* (3B), \*\*Tiny\*\* (7B/1B actifs) et \*\*Small\*\* (32B/9B actifs). Entraînée sur 15T de jetons, la nouvelle architecture hybride (H) Mamba d’IBM permet aux modèles Granite-4.0 de fonctionner plus rapidement avec une utilisation mémoire réduite. Découvrez \[comment exécuter\](#run-granite-4.0-tutorials) Unsloth Granite-4.0 Dynamic GGUF ou affiner/RL le modèle. Vous pouvez \[affiner Granite-4.0\](#fine-tuning-granite-4.0-in-unsloth) avec notre notebook Colab gratuit pour un cas d’utilisation d’agent de support. [Tutoriel d’exécution](https://unsloth.ai/docs/fr/modeles/tutorials/ibm-granite-4.0.md#run-granite-4.0-tutorials) [Tutoriel de fine-tuning](https://unsloth.ai/docs/fr/modeles/tutorials/ibm-granite-4.0.md#fine-tuning-granite-4.0-in-unsloth) \*\*Téléchargements Unsloth Granite-4.0 :\*\* | Dynamic GGUF | Dynamic 4-bit + FP8 | Instruct 16 bits | | --- | --- | --- | | * [H-350M](https://huggingface.co/unsloth/granite-4.0-h-350m-GGUF)

* [350M](https://huggingface.co/unsloth/granite-4.0-350m-GGUF)

* [H-1B](https://huggingface.co/unsloth/granite-4.0-h-1b-GGUF)

* [1B](https://huggingface.co/unsloth/granite-4.0-1b-GGUF)

* [H-Small](https://huggingface.co/unsloth/granite-4.0-h-small-GGUF)

* [H-Tiny](https://huggingface.co/unsloth/granite-4.0-h-tiny-GGUF)

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro-GGUF)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro-GGUF) | Dynamic 4-bit Instruct :

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro-unsloth-bnb-4bit)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro-unsloth-bnb-4bit)


FP8 Dynamic :

* [H-Small FP8](https://huggingface.co/unsloth/granite-4.0-h-small-FP8-Dynamic)

* [H-Tiny FP8](https://huggingface.co/unsloth/granite-4.0-h-tiny-FP8-Dynamic) | * [H-350M](https://huggingface.co/unsloth/granite-4.0-h-350m)

* [350M](https://huggingface.co/unsloth/granite-4.0-350m)

* [H-1B](https://huggingface.co/unsloth/granite-4.0-h-1b)

* [1B](https://huggingface.co/unsloth/granite-4.0-1b)

* [H-Small](https://huggingface.co/unsloth/granite-4.0-h-small)

* [H-Tiny](https://huggingface.co/unsloth/granite-4.0-h-tiny)

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro) | Vous pouvez également consulter notre \[collection Granite-4.0\](https://huggingface.co/collections/unsloth/granite-40-68ddf64b4a8717dc22a9322d) pour tous les fichiers téléversés, y compris les quants Dynamic Float8, etc. \*\*Explications des modèles Granite-4.0 :\*\* \* \*\*Nano et H-Nano :\*\* Les modèles 350M et 1B offrent de solides capacités de suivi d’instructions, permettant des applications avancées d’IA embarquée et en périphérie, ainsi que de recherche/fine-tuning. \* \*\*H-Small (MoE) :\*\* Cheval de bataille d’entreprise pour les tâches quotidiennes, prend en charge plusieurs sessions à long contexte sur des GPU d’entrée de gamme comme le L40S (32B au total, 9B actifs). \* \*\*H-Tiny (MoE) :\*\* Rapide, économique pour les tâches à grand volume et faible complexité ; optimisé pour une utilisation locale et en périphérie (7B au total, 1B actif). \* \*\*H-Micro (Dense) :\*\* Léger, efficace pour les charges de travail à grand volume et faible complexité ; idéal pour un déploiement local et en périphérie (3B au total). \* \*\*Micro (Dense) :\*\* Option dense alternative lorsque Mamba2 n’est pas entièrement pris en charge (3B au total). ## Exécuter les tutoriels Granite-4.0 ### :gear: Paramètres d’inférence recommandés IBM recommande ces paramètres : \`temperature=0.0\`, \`top\_p=1.0\`, \`top\_k=0\` \* \*\*Température de 0.0\*\* \* Top\\\_K = 0 \* Top\\\_P = 1.0 \* Contexte minimum recommandé : 16 384 \* Fenêtre de longueur de contexte maximale : 131 072 (contexte 128K) \*\*Modèle de chat :\*\* \`\`\` <|start\_of\_role|>system<|end\_of\_role|>Vous êtes un assistant utile. Veuillez vous assurer que les réponses sont professionnelles, exactes et sûres.<|end\_of\_text|> <|start\_of\_role|>user<|end\_of\_role|>Veuillez citer un laboratoire IBM Research situé aux États-Unis. Vous devez uniquement afficher son nom et son emplacement.<|end\_of\_text|> <|start\_of\_role|>assistant<|end\_of\_role|>Almaden Research Center, San Jose, California<|end\_of\_text|> \`\`\` ### :llama: Ollama : tutoriel pour exécuter Granite-4.0 1. Installez \`ollama\` si ce n’est pas déjà fait ! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Exécutez le modèle ! Notez que vous pouvez appeler \`ollama serve\`dans un autre terminal si cela échoue ! Nous incluons toutes nos corrections et les paramètres suggérés (température, etc.) dans \`params\` dans notre téléversement sur Hugging Face ! Vous pouvez changer le nom du modèle '\`granite-4.0-h-small-GGUF\`' en n’importe quel modèle Granite comme 'granite-4.0-h-micro:Q8\\\_K\\\_XL'. \`\`\`bash ollama run hf.co/unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` ### 📖 llama.cpp : tutoriel pour exécuter Granite-4.0 1. Obtenez la dernière version de \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Remplacez \`-DGGML\_CUDA=ON\` par \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d’habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Si vous souhaitez utiliser \`llama.cpp\` directement pour charger des modèles, vous pouvez faire ce qui suit : (:Q4\\\_K\\\_XL) est le type de quantification. Vous pouvez également télécharger via Hugging Face (point 3). C’est similaire à \`ollama run\` \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` 3. \*\*OU\*\* téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir Q4\\\_K\\\_M, ou d’autres versions quantifiées (comme BF16 en pleine précision). \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/granite-4.0-h-small-GGUF", local\_dir = "unsloth/granite-4.0-h-small-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], # Pour Q4\_K\_M ) \`\`\` 4. Exécutez le test Flappy Bird d’Unsloth 5. Modifiez \`--threads 32\` pour le nombre de threads CPU, \`--ctx-size 16384\` pour la longueur de contexte (Granite-4.0 prend en charge une longueur de contexte de 128K !), \`--n-gpu-layers 99\` pour le déchargement GPU, selon le nombre de couches. Essayez de l’ajuster si votre GPU manque de mémoire. Supprimez-le également si vous n’avez qu’une inférence CPU. 6. Pour le mode conversation : \`\`\`bash ./llama.cpp/llama-mtmd-cli \\ --model unsloth/granite-4.0-h-small-GGUF/granite-4.0-h-small-UD-Q4\_K\_XL.gguf \\ --jinja \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 0.0 \\ --top-k 0 \\ --top-p 1.0 \`\`\` ### 🐋 Docker : tutoriel pour exécuter Granite-4.0 Si vous avez déjà Docker Desktop, tout ce que vous avez à faire est d’exécuter la commande ci-dessous et c’est terminé : \`\`\`bash docker model pull hf.co/unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` ## :sloth: Fine-tuning de Granite-4.0 dans Unsloth Unsloth prend désormais en charge tous les modèles Granite 4.0, y compris nano, micro, tiny et small pour le fine-tuning. L’entraînement est 2x plus rapide, utilise 50 % de VRAM en moins et prend en charge des longueurs de contexte 6x plus longues. Granite-4.0 micro et tiny tiennent confortablement dans un GPU T4 de 15 Go de VRAM. \* \*\*Granite-4.0\*\* \[\*\*notebook gratuit de fine-tuning\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Granite4.0.ipynb) \* Granite-4.0-350M \[notebook de fine-tuning\](https://github.com/unslothai/notebooks/blob/main/nb/Granite4.0\_350M.ipynb) Ce notebook entraîne un modèle pour devenir un agent de support qui comprend les interactions clients, avec analyses et recommandations. Cette configuration vous permet d’entraîner un bot qui fournit une assistance en temps réel aux agents de support. Nous vous montrons également comment entraîner un modèle à l’aide de données stockées dans une feuille Google. ![](https://unsloth.ai/files/a8a63aa9f1d816451ec6813e065b66fac12e5e16) \*\*Configuration Unsloth pour Granite-4.0 :\*\* \`\`\`python !pip install --upgrade unsloth from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/granite-4.0-h-micro", max\_seq\_length = 2048, # Longueur de contexte - peut être plus longue, mais utilise plus de mémoire load\_in\_4bit = True, # Le 4 bits utilise beaucoup moins de mémoire load\_in\_8bit = False, # Un peu plus précis, utilise 2x plus de mémoire full\_finetuning = False, # Nous avons maintenant le fine-tuning complet ! # token = "hf\_...", # utilisez-en un si vous utilisez des modèles à accès restreint ) \`\`\` Si vous avez une ancienne version d’Unsloth et/ou si vous effectuez le fine-tuning en local, installez la dernière version d’Unsloth : \`\`\` pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/ibm-granite-4.0.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/ministral-3.md). # Ministral 3 - Guide d'exécution Mistral lance Ministral 3, ses nouveaux modèles multimodaux en variantes Base, Instruct et Reasoning, disponibles en \*\*3B\*\*, \*\*8B\*\*, et \*\*14B\*\* tailles. Ils offrent des performances de pointe pour leur taille et sont affinés pour les cas d’usage d’instructions et de chat. Les modèles multimodaux prennent en charge \*\*contexte de 256K\*\* les fenêtres, plusieurs langues, l’appel de fonctions natif et la sortie JSON. Le modèle complet Ministral-3-Instruct-2512 non quantifié de 14B tient dans \*\*24 Go de RAM\*\*/VRAM. Vous pouvez désormais exécuter, affiner et faire du RL sur tous les modèles Ministral 3 avec Unsloth : [Exécuter les tutoriels Ministral 3](https://unsloth.ai/docs/fr/modeles/tutorials/ministral-3.md#run-ministral-3-tutorials) [Affinage de Ministral 3](https://unsloth.ai/pages/b7e8e6530edc0f8d90d1dd7614a8a127f98532da#fine-tuning) Nous avons également téléversé Mistral Large 3 \[GGUF ici\](https://huggingface.co/unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF). Pour tous les téléversements Ministral 3 (BnB, FP8), \[voir ici\](https://huggingface.co/collections/unsloth/ministral-3). | GGUF de Ministral-3-Instruct : | GGUF de Ministral-3-Reasoning : | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | \[3B\](https://huggingface.co/unsloth/Ministral-3-3B-Instruct-2512-GGUF) • \[8B\](https://huggingface.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF) • \[14B\](https://huggingface.co/unsloth/Ministral-3-14B-Instruct-2512-GGUF) | \[3B\](https://huggingface.co/unsloth/Ministral-3-3B-Reasoning-2512-GGUF) • \[8B\](https://huggingface.co/unsloth/Ministral-3-8B-Reasoning-2512-GGUF) • \[14B\](https://huggingface.co/unsloth/Ministral-3-14B-Reasoning-2512-GGUF) | ### ⚙️ Guide d’utilisation Pour obtenir des performances optimales pour \*\*Instruct\*\*, Mistral recommande d’utiliser des températures plus basses telles que \`temperature = 0.15\` ou \`0.1\` Pour \*\*Reasoning\*\*, Mistral recommande \`température = 0.7\` et \`top\_p = 0.95\`. | Instruct : | Reasoning : | | ----------------------------- | ------------------- | | \`Température = 0,15\` ou \`0.1\` | \`Température = 0.7\` | | \`Top\_P = par défaut\` | \`Top\_P = 0.95\` | \*\*Longueur de sortie adéquate\*\*: Utilisez une longueur de sortie de \`32,768\` tokens pour la plupart des requêtes pour la variante reasoning, et \`16,384\` pour la variante instruct. Vous pouvez augmenter la taille maximale de sortie pour le modèle reasoning si nécessaire. La longueur maximale du contexte que Ministral 3 peut atteindre est \`262,144\` Le format du modèle de conversation se trouve lorsque nous utilisons ce qui suit : {% code overflow="wrap" %} \`\`\`python tokenizer.apply\_chat\_template(\[ {"role" : "user", "content" : "What is 1+1?"}, {"role" : "assistant", "content" : "2"}, {"role" : "user", "content" : "What is 2+2?"} \], add\_generation\_prompt = True ) \`\`\` {% endcode %} #### Ministral \*Reasoning\* template de chat : {% code overflow="wrap" lineNumbers="true" %} \`\`\` ~\[SYSTEM\_PROMPT\]# COMMENT VOUS DEVEZ RÉFLÉCHIR ET RÉPONDRE Rédigez d’abord votre processus de réflexion (monologue intérieur) jusqu’à obtenir une réponse. Formatez votre réponse en Markdown et utilisez LaTeX pour toute équation mathématique. Écrivez à la fois vos pensées et la réponse dans la même langue que l’entrée. Votre processus de réflexion doit suivre le modèle ci-dessous :\[THINK\]Vos pensées et/ou brouillon, comme si vous travailliez sur un exercice sur une feuille de brouillon. Soyez aussi informel et aussi long que vous le souhaitez jusqu’à ce que vous soyez certain de générer la réponse à l’utilisateur.\[/THINK\]Ici, fournissez une réponse autonome.\[/SYSTEM\_PROMPT\]\[INST\]Que fait 1+1 ?\[/INST\]2~\[INST\]Que fait 2+2 ?\[/INST\] \`\`\` {% endcode %} #### Ministral \*Instruct\* template de chat : {% code overflow="wrap" lineNumbers="true" expandable="true" %} \`\`\` ~\[SYSTEM\_PROMPT\]Vous êtes Ministral-3-3B-Instruct-2512, un grand modèle de langage (LLM) créé par Mistral AI, une startup française dont le siège est à Paris. Vous alimentez un assistant IA appelé Le Chat. Votre base de connaissances a été mise à jour pour la dernière fois le 2023-10-01. La date actuelle est {today}. Lorsque vous n’êtes pas sûr de certaines informations ou lorsque la demande de l’utilisateur nécessite des données à jour ou spécifiques, vous devez utiliser les outils disponibles pour récupérer les informations. N’hésitez pas à utiliser les outils chaque fois qu’ils peuvent fournir une réponse plus précise ou plus complète. Si aucun outil pertinent n’est disponible, indiquez clairement que vous ne disposez pas de l’information et évitez d’inventer quoi que ce soit. Si la question de l’utilisateur n’est pas claire, est ambiguë ou ne fournit pas suffisamment de contexte pour que vous puissiez y répondre avec précision, n’essayez pas d’y répondre immédiatement et demandez plutôt à l’utilisateur de clarifier sa demande (par ex. "Quels sont de bons restaurants près de chez moi ?" => "Où êtes-vous ?" ou "Quand est le prochain vol pour Tokyo" => "D’où partez-vous ?"). Vous faites toujours très attention aux dates, en particulier vous essayez de résoudre les dates (par ex. "hier" correspond à {yesterday}) et, lorsqu’on vous demande des informations à des dates spécifiques, vous ignorez les informations qui concernent une autre date. Vous suivez ces instructions dans toutes les langues et répondez toujours à l’utilisateur dans la langue qu’il utilise ou demande. Les sections suivantes décrivent les capacités dont vous disposez. # INSTRUCTIONS DE NAVIGATION WEB Vous ne pouvez effectuer aucune recherche web ni accéder à Internet pour ouvrir des URL, des liens, etc. Si cela semble être ce que l’utilisateur attend de vous, vous clarifiez la situation et demandez à l’utilisateur de copier-coller le texte directement dans le chat. # INSTRUCTIONS MULTIMODALES Vous avez la capacité de lire des images, mais vous ne pouvez pas générer d’images. Vous ne pouvez pas non plus transcrire des fichiers audio ou des vidéos. Vous ne pouvez ni lire ni transcrire des fichiers audio ou des vidéos. # INSTRUCTIONS D’APPEL D’OUTILS Vous pouvez avoir accès à des outils que vous pouvez utiliser pour récupérer des informations ou effectuer des actions. Vous devez utiliser ces outils dans les situations suivantes : 1. Lorsque la demande nécessite des informations à jour. 2. Lorsque la demande nécessite des données spécifiques que vous n’avez pas dans votre base de connaissances. 3. Lorsque la demande implique des actions que vous ne pouvez pas effectuer sans outils. Priorisez toujours l’utilisation des outils afin de fournir la réponse la plus précise et la plus utile. Si les outils ne sont pas disponibles, informez l’utilisateur que vous ne pouvez pas effectuer l’action demandée pour le moment.\[/SYSTEM\_PROMPT\]\[INST\]Que fait 1+1 ?\[/INST\]2~\[INST\]Que fait 2+2 ?\[/INST\] \`\`\` {% endcode %} ## 📖 Exécuter les tutoriels Ministral 3 Voici des guides pour les \[Reasoning\](#reasoning-ministral-3-reasoning-2512) et \[Instruct\](#instruct-ministral-3-instruct-2512) variantes du modèle. ### Instruct : Ministral-3-Instruct-2512 Pour obtenir des performances optimales pour \*\*Instruct\*\*, Mistral recommande d’utiliser des températures plus basses telles que \`temperature = 0.15\` ou \`0.1\` #### :sparkles: Llama.cpp : Exécuter le tutoriel Ministral-3-14B-Instruct {% stepper %} {% step %} Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} Vous pouvez le récupérer directement depuis Hugging Face via : \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Ministral-3-14B-Instruct-2512-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32784 \\ --temp 0,15 \`\`\` {% endstep %} {% step %} Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir \`UD\_Q4\_K\_XL\` ou d’autres versions quantifiées. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Ministral-3-14B-Instruct-2512-GGUF", local\_dir = "Ministral-3-14B-Instruct-2512-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` {% endstep %} {% endstepper %} ### Reasoning : Ministral-3-Reasoning-2512 Pour obtenir des performances optimales pour \*\*Reasoning\*\*, Mistral recommande d’utiliser \`température = 0.7\` et \`top\_p = 0.95\`. #### :sparkles: Llama.cpp : Exécuter le tutoriel Ministral-3-14B-Reasoning {% stepper %} {% step %} Obtenez la dernière version \`llama.cpp\` sur \[GitHub\](https://github.com/ggml-org/llama.cpp). Vous pouvez également utiliser les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} Vous pouvez le récupérer directement depuis Hugging Face via : \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Ministral-3-14B-Reasoning-2512-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32784 \\ --temp 0.6 --top-p 0.95 \`\`\` {% endstep %} {% step %} Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir \`UD\_Q4\_K\_XL\` ou d’autres versions quantifiées. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Ministral-3-14B-Reasoning-2512-GGUF", local\_dir = "Ministral-3-14B-Reasoning-2512-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` {% endstep %} {% endstepper %} ## 🛠️ Affinage de Ministral 3 [](https://unsloth.ai/docs/fr/modeles/tutorials/ministral-3.md#fine-tuning) Unsloth prend désormais en charge l’affinage de tous les modèles Ministral 3, y compris la prise en charge de la vision. Pour entraîner, vous devez utiliser la dernière version de 🤗Hugging Face \`transformers\` v5 et \`unsloth\` qui inclut notre récente \[très long contexte\](/docs/fr/blog/500k-context-length-fine-tuning.md) prise en charge. Le grand modèle Ministral 3 de 14B devrait tenir sur un GPU Colab gratuit. Nous avons créé des notebooks Unsloth gratuits pour affiner Ministral 3. Modifiez le nom pour utiliser le modèle souhaité. \* Ministral-3B-Instruct \[Notebook vision\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_VL\_\\(3B\\)\_Vision.ipynb) (vision) \* Ministral-3B-Instruct \[Notebook GRPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) {% columns %} {% column %} Notebook d’affinage de Ministral Vision {% embed url="" %} {% endcolumn %} {% column %} Notebook GRPO RL de Ministral Sudoku {% embed url="" %} {% endcolumn %} {% endcolumns %} ### :sparkles:Apprentissage par renforcement (GRPO) Unsloth prend désormais également en charge le RL et le GRPO pour les modèles Mistral. Comme d’habitude, ils bénéficient de toutes les améliorations d’Unsloth et, demain, nous allons bientôt publier un notebook spécialement destiné à résoudre automatiquement le Sudoku. \* Ministral-3B-Instruct \[Notebook GRPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) \*\*Pour utiliser la dernière version d’Unsloth et de transformers v5, mettez à jour via :\*\* {% code overflow="wrap" %} \`\`\` pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth\_zoo \`\`\` {% endcode %} L’objectif est de générer automatiquement des stratégies pour résoudre le Sudoku ! {% columns %} {% column %} ![](https://unsloth.ai/files/26269a7835a9fcfac58fb39f8050c9db8426b4af) {% endcolumn %} {% column %} ![](https://unsloth.ai/files/8e81bf759f94c0b2a94f6592d8de036589765203) {% endcolumn %} {% endcolumns %} Pour les graphiques de récompense pour Ministral, nous obtenons ce qui suit. On voit que cela fonctionne bien ! {% columns %} {% column %} !\[\](/files/759e8aee0bebc233f95e575755d043cd5fbb3574) !\[\](/files/3abede18b3903d0b4c897967b7f1c26baee12ba9) {% endcolumn %} {% column %} !\[\](/files/04333c4fdd01411fc0c2b7f09bab41f949397d4c) !\[\](/files/e25f77c1e2d627e5e3541bbc612ab0073f4079e0) {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/ministral-3.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/cogito-v2-how-to-run-locally.md). # Cogito v2.1: How to Run Locally {% hint style="success" %} Deep Cogito v2.1 is an updated 671B MoE that is the most powerful open weights model as of 19 November 2025. {% endhint %} Cogito v2.1 comes in 1 671B MoE size, whilst Cogito v2 Preview is \[Deep Cogito\](https://www.deepcogito.com/)'s release of models spans 4 model sizes ranging from 70B to 671B. By using \*\*IDA (Iterated Distillation & Amplification)\*\*, these models are trained with the model internalizing the reasoning process using iterative policy improvement, rather than simply searching longer at inference time (like DeepSeek R1). Deep Cogito is based in \[San Fransisco, USA\](https://techcrunch.com/2025/04/08/deep-cogito-emerges-from-stealth-with-hybrid-ai-reasoning-models/) (like Unsloth :flag\\\_us:) and we're excited to provide quantized dynamic models for all 4 model sizes! All uploads use Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) for SOTA 5-shot MMLU and KL Divergence performance, meaning you can run & fine-tune quantized these LLMs with minimal accuracy loss! \*\*Tutorials navigation:\*\* [Run 671B MoE](https://docs.unsloth.ai/basics/tutorials-how-to-fine-tune-and-run-llms/cogito-v2-how-to-run-locally#run-cogito-671b-moe-in-llama.cpp) [Run 109B MoE](https://docs.unsloth.ai/basics/tutorials-how-to-fine-tune-and-run-llms/cogito-v2-how-to-run-locally#run-cogito-109b-moe-in-llama.cpp) [Run 405B Dense](https://docs.unsloth.ai/basics/tutorials-how-to-fine-tune-and-run-llms/cogito-v2-how-to-run-locally#run-cogito-405b-dense-in-llama.cpp) [Run 70B Dense](https://docs.unsloth.ai/basics/tutorials-how-to-fine-tune-and-run-llms/cogito-v2-how-to-run-locally#run-cogito-70b-dense-in-llama.cpp) {% hint style="success" %} Choose which model size fits your hardware! We upload 1.58bit to 16bit variants for all 4 model sizes! {% endhint %} ## :gem: Model Sizes and Uploads There are 4 model sizes: 1. 2 Dense models based off from Llama - 70B and 405B 2. 2 MoE models based off from Llama 4 Scout (109B) and DeepSeek R1 (671B) | Model Sizes | Recommended Quant & Link | Disk Size | Architecture | | --- | --- | --- | --- | | 70B Dense | [UD-Q4\_K\_XL](https://huggingface.co/unsloth/cogito-v2-preview-llama-70B-GGUF) | **44GB** | Llama 3 70B | | 109B MoE | [UD-Q3\_K\_XL](https://huggingface.co/unsloth/cogito-v2-preview-llama-109B-MoE-GGUF) | **50GB** | Llama 4 Scout | | 405B Dense | [UD-Q2\_K\_XL](https://huggingface.co/unsloth/cogito-v2-preview-llama-405B-GGUF) | **152GB** | Llama 3 405B | | 671B MoE | [UD-Q2\_K\_XL](https://huggingface.co/unsloth/cogito-v2-preview-deepseek-671B-MoE-GGUF) | **251GB** | DeepSeek R1 | {% hint style="success" %} Though not necessary, for the best performance, have your VRAM + RAM combined = to the size of the quant you're downloading. If you have less VRAM + RAM, then the quant will still function, just be much slower. {% endhint %} ## 🐳 Run Cogito 671B MoE in llama.cpp 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. {% code overflow="wrap" %} \`\`\`shellscript apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} 2. If you want to use \`llama.cpp\` directly to load models, you can do the below: (:IQ1\\\_S) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. {% hint style="success" %} Please try out \`-ot ".ffn\_.\*\_exps.=CPU"\` to offload all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. And finally offload all layers via \`-ot ".ffn\_.\*\_exps.=CPU"\` This uses the least VRAM. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. {% endhint %} \`\`\`shellscript export LLAMA\_CACHE="unsloth/cogito-671b-v2.1-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/cogito-671b-v2.1-GGUF:UD-Q2\_K\_XL \\ --n-gpu-layers 99 \\ --temp 0.6 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ --jinja \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` 3. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD-IQ1\_S\`(dynamic 1.78bit quant) or other quantized versions like \`Q4\_K\_M\` . We \*\*recommend using our 2.7bit dynamic quant\*\*\*\* \*\*\*\*\`UD-Q2\_K\_XL\`\*\*\*\* \*\*\*\*to balance size and accuracy\*\*. More versions at: {% code overflow="wrap" %} \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "0" # Can sometimes rate limit, so set to 0 to disable from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/cogito-671b-v2.1-GGUF", local\_dir = "unsloth/cogito-671b-v2.1-GGUF", allow\_patterns = \["\*UD-IQ1\_S\*"\], # Dynamic 1bit (168GB) Use "\*UD-Q2\_K\_XL\*" for Dynamic 2bit (251GB) ) \`\`\` {% endcode %} 4. Edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length, \`--n-gpu-layers 2\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. ## :mouse\\\_three\\\_button:Run Cogito 109B MoE in llama.cpp 1. Follow the same instructions as running the \[671B model above\](#run-cogito-671b-moe-in-llama.cpp). 2. Then run the below: \`\`\`shellscript export LLAMA\_CACHE="unsloth/cogito-v2-preview-llama-109B-MoE-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/cogito-v2-preview-llama-109B-MoE-GGUF:Q3\_K\_XL \\ --n-gpu-layers 99 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ --ctx-size 16384 \\ --jinja \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` ## :deciduous\\\_tree:Run Cogito 405B Dense in llama.cpp 1. Follow the same instructions as running the \[671B model above\](#run-cogito-671b-moe-in-llama.cpp). 2. Then run the below: \`\`\`shellscript export LLAMA\_CACHE="unsloth/cogito-v2-preview-llama-405B-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/cogito-v2-preview-llama-405B-GGUF:Q2\_K\_XL \\ --n-gpu-layers 99 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ --jinja \\ --ctx-size 16384 \`\`\` ## :sunglasses: Run Cogito 70B Dense in llama.cpp 1. Follow the same instructions as running the \[671B model above\](#run-cogito-671b-moe-in-llama.cpp). 2. Then run the below: \`\`\`shellscript export LLAMA\_CACHE="unsloth/cogito-v2-preview-llama-70B-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/cogito-v2-preview-llama-70B-GGUF:Q4\_K\_XL \\ --n-gpu-layers 99 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ --jinja \\ --ctx-size 16384 \`\`\` See for more details --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/cogito-v2-how-to-run-locally.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/qwen-image-2512/stable-diffusion.cpp.md). # Run Qwen-Image-2512 in stable-diffusion.cpp Tutorial \*\*Qwen-Image-2512\*\* is Qwen's new text-to-image foundational model and you can now run it on your local device via stable-diffusion.cpp. See below for instructions: ## 📖 stable-diffusion.cpp Tutorial \[stable-diffusion.cpp\](https://github.com/leejet/stable-diffusion.cpp) is an open-source library for efficient and local inference of diffusion image models written in pure C/C++. To run, you don't need a GPU, just a CPU with RAM will work. For best results, ensure your total usable memory (RAM + VRAM / unified) is larger than the GGUF size; e.g. 4-bit (Q4\\\_K\\\_M) \`unsloth/Qwen-Image-Edit-2512-GGUF\` is 13.1 GB, so you should have 13.2+ GB of combined memory. The tutorial will focus on machines with CUDA available, but instructions to build with on Apple or CPU only are similar and available in the repo. ### #1. Setup environment We will be building from source so we need to first be sure your build software is installed \`\`\`bash sudo apt update sudo apt install -y git cmake build-essential pkg-config \`\`\` {% hint style="info" %} \[Releases Page\](https://github.com/leejet/stable-diffusion.cpp/releases) may have pre built binaries available for your hardware if you don't want to go through the build process. {% endhint %} Make sure CUDA environment variables are set: \`\`\`bash export CUDA\_HOME=/usr/local/cuda export PATH="$CUDA\_HOME/bin:$PATH" export LD\_LIBRARY\_PATH="$CUDA\_HOME/lib64:${LD\_LIBRARY\_PATH:-}" \`\`\` You can confirm if set correctly by running: \`\`\`bash nvcc --version // if not found install nvidia-cuda-toolkit ldconfig -p | grep -E 'libcudart\\.so|libcublas\\.so' \`\`\` We can now clone the repo and build: {% code expandable="true" %} \`\`\`bash git clone --recursive https://github.com/leejet/stable-diffusion.cpp cd stable-diffusion.cpp mkdir -p build cd build cmake .. -DCMAKE\_BUILD\_TYPE=Release -DSD\_CUDA=ON cmake --build . -j"$(nproc)" \`\`\` {% endcode %} Confirm sd-cli was built: \`\`\`bash ls bin/sd-cli \`\`\` ### #2. Download Models Diffusion models typically need 3 components. A Variational AutoEncoder (VAE) that encodes image pixel space to latent space, a text encoder to translate text to input embeddings, and the actual diffusion transformer. Both the diffusion model and text encoder can be GGUF format while we typically use safetensors for the vae. Let's download the models we will use: \`\`\`bash cd .. mkdir models mkdir outputs ## Diffusion Models curl -L -C - -o models/qwen-image-2512-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-2512-GGUF/resolve/main/qwen-image-2512-Q4\_K\_M.gguf curl -L -C - -o models/qwen-image-edit-2511-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF/resolve/main/qwen-image-edit-2511-Q4\_K\_M.gguf ## Text Encoder + VAE curl -L -C - -o models/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf \\ https://huggingface.co/unsloth/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf curl -L -C - -o models/qwen\_image\_vae.safetensors \\ https://huggingface.co/Comfy-Org/Qwen-Image\_ComfyUI/resolve/main/split\_files/vae/qwen\_image\_vae.safetensors \`\`\` We are using Q4 GGUF variants, but you can try smaller or larger quant types depending on how much VRAM/RAM you have. {% hint style="warning" %} The format of the vae and diffusion model might be different than the diffusers checkpoints. Only use checkpoints that are compatible with stable-diffusion.cpp and ComfyUI. {% endhint %} #### Workflow and Hyperparameters You can view our detailed \[Run GGUFs in ComfyUI\](/docs/blog/comfyui.md#workflow-and-hyperparameters-1) Guide. ### #3. Inference We can now run the binary that we built. This is an example of a basic text to image command: \`\`\`bash ./build/bin/sd-cli --diffusion-model models/qwen-image-2512-Q4\_K\_M.gguf \\ --vae models/qwen\_image\_vae.safetensors \\ --llm models/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf \\ --cfg-scale 2.5 --sampling-method euler -v --steps 40 \\ -H 1024 -W 1024 --diffusion-fa --flow-shift 3 \\ -p 'Aerial drone photograph of a vast field of bright yellow wildflowers with the text "Unsloth + Diffusion" spelled out in deep purple lavender flowers, sharp contrast between yellow and purple, natural organic letter shapes formed by flower beds, golden hour lighting, rolling countryside landscape, high altitude perspective looking straight down, photorealistic, 8K resolution' \\ --offload-to-cpu -o outputs/unsloth\_diffusion.png \`\`\` {% hint style="success" %} No need for \`--offload-to-cpu\` if you have enough VRAM. {% endhint %} ![](https://unsloth.ai/files/OjWwnhucl82TwBpUJMAQ) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/qwen-image-2512/stable-diffusion.cpp.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md). # DeepSeek-R1-0528: How to Run Locally DeepSeek-R1-0528 is DeepSeek's new update to their R1 reasoning model. The full 671B parameter model requires 715GB of disk space. The quantized dynamic \*\*1.66-bit\*\* version uses 162GB (-80% reduction in size). GGUF: \[DeepSeek-R1-0528-GGUF\](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF) DeepSeek also released a R1-0528 distilled version by fine-tuning Qwen3 (8B). The distill achieves similar performance to Qwen3 (235B). \*\*\*You can also\*\*\* \[\*\*\*fine-tune Qwen3 Distill\*\*\*\](#fine-tuning-deepseek-r1-0528-with-unsloth) \*\*\*with Unsloth\*\*\*. Qwen3 GGUF: \[DeepSeek-R1-0528-Qwen3-8B-GGUF\](https://huggingface.co/unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF) All uploads use Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) for SOTA 5-shot MMLU and KL Divergence performance, meaning you can run & fine-tune quantized DeepSeek LLMs with minimal accuracy loss. \*\*Tutorials navigation:\*\* [Run in llama.cpp](https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md#run-qwen3-distilled-r1-in-llama.cpp) [Run in Ollama/Open WebUI](https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md#run-in-ollama-open-webui) [Fine-tuning R1-0528](https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md#fine-tuning-deepseek-r1-0528-with-unsloth) {% hint style="success" %} NEW: Huge improvements to tool calling and chat template fixes.\\ \\ New \[TQ1\\\_0 dynamic 1.66-bit quant\](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF?show\_file\_info=DeepSeek-R1-0528-UD-TQ1\_0.gguf) - 162GB in size. Ideal for 192GB RAM (including Mac) and Ollama users. Try: \`ollama run hf.co/unsloth/DeepSeek-R1-0528-GGUF:TQ1\_0\` {% endhint %} ## :gear: Recommended Settings For DeepSeek-R1-0528-Qwen3-8B, the model can pretty much fit in any setup, and even those with as less as 20GB RAM. There is no need for any prep beforehand.\\ \\ However, for the full R1-0528 model which is 715GB in size, you will need extra prep. The 1.78-bit (IQ1\\\_S) quant will fit in a 1x 24GB GPU (with all layers offloaded). Expect around 5 tokens/s with this setup if you have bonus 128GB RAM as well. It is recommended to have at least 64GB RAM to run this quant (you will get 1 token/s without a GPU). For optimal performance you will need at least \*\*180GB unified memory or 180GB combined RAM+VRAM\*\* for 5+ tokens/s. We suggest using our 2.7bit (Q2\\\_K\\\_XL) or 2.4bit (IQ2\\\_XXS) quant to balance size and accuracy! The 2.4bit one also works well. {% hint style="success" %} Though not necessary, for the best performance, have your VRAM + RAM combined = to the size of the quant you're downloading. {% endhint %} ### 🐳 Official Recommended Settings: According to \[DeepSeek\](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528), these are the recommended settings for R1 (R1-0528 and Qwen3 distill should use the same settings) inference: \* Set the \*\*temperature 0.6\*\* to reduce repetition and incoherence. \* Set \*\*top\\\_p to 0.95\*\* (recommended) \* Run multiple tests and average results for reliable evaluation. ### :1234: Chat template/prompt format R1-0528 uses the same chat template as the original R1 model. You do not need to force \`\\n\` , but you can still add it in! \`\`\` <|begin▁of▁sentence|><|User|>What is 1+1?<|Assistant|>It's 2.<|end▁of▁sentence|><|User|>Explain more!<|Assistant|> \`\`\` A BOS is forcibly added, and an EOS separates each interaction. To counteract double BOS tokens during inference, you should only call \`tokenizer.encode(..., add\_special\_tokens = False)\` since the chat template auto adds a BOS token as well.\\ For llama.cpp / GGUF inference, you should skip the BOS since it’ll auto add it: \`\`\` <|User|>What is 1+1?<|Assistant|> \`\`\` The \`\` and \`\` tokens get their own designated tokens. ## Model uploads \*\*ALL our uploads\*\* - including those that are not imatrix-based or dynamic, utilize our calibration dataset, which is specifically optimized for conversational, coding, and language tasks. \* Qwen3 (8B) distill: \[DeepSeek-R1-0528-Qwen3-8B-GGUF\](https://huggingface.co/unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF) \* Full DeepSeek-R1-0528 model uploads below: We also uploaded \[IQ4\\\_NL\](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/IQ4\_NL) and \[Q4\\\_1\](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/Q4\_1) quants which run specifically faster for ARM and Apple devices respectively. | MoE Bits | Type + Link | Disk Size | Details | | --- | --- | --- | --- | | 1.66bit | [TQ1\_0](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF?show_file_info=DeepSeek-R1-0528-UD-TQ1_0.gguf) | **162GB** | 1.92/1.56bit | | 1.78bit | [IQ1\_S](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-IQ1_S) | **185GB** | 2.06/1.56bit | | 1.93bit | [IQ1\_M](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-IQ1_M) | **200GB** | 2.5/2.06/1.56 | | 2.42bit | [IQ2\_XXS](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-IQ2_XXS) | **216GB** | 2.5/2.06bit | | 2.71bit | [Q2\_K\_XL](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-Q2_K_XL) | **251GB** | 3.5/2.5bit | | 3.12bit | [IQ3\_XXS](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-IQ3_XXS) | **273GB** | 3.5/2.06bit | | 3.5bit | [Q3\_K\_XL](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-Q3_K_XL) | **296GB** | 4.5/3.5bit | | 4.5bit | [Q4\_K\_XL](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-Q4_K_XL) | **384GB** | 5.5/4.5bit | | 5.5bit | [Q5\_K\_XL](https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF/tree/main/UD-Q5_K_XL) | **481GB** | 6.5/5.5bit | We've also uploaded versions in \[BF16 format\](https://huggingface.co/unsloth/DeepSeek-R1-0528-BF16), and original \[FP8 (float8) format\](https://huggingface.co/unsloth/DeepSeek-R1-0528). ## Run DeepSeek-R1-0528 Tutorials: ### :llama: Run in Ollama/Open WebUI 1. Install \`ollama\` if you haven't already! You can only run models up to 32B in size. To run the full 720GB R1-0528 model, \[see here\](#run-full-r1-0528-on-ollama-open-webui). \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Run the model! Note you can call \`ollama serve\`in another terminal if it fails! We include all our fixes and suggested parameters (temperature etc) in \`params\` in our Hugging Face upload! \`\`\`bash ollama run hf.co/unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF:Q4\_K\_XL \`\`\` 3. \*\*(NEW) To run the full R1-0528 model in Ollama, you can use our TQ1\\\_0 (162GB quant):\*\* \`\`\`bash OLLAMA\_MODELS=unsloth\_downloaded\_models ollama serve & ollama run hf.co/unsloth/DeepSeek-R1-0528-GGUF:TQ1\_0 \`\`\` ### :llama: Run Full R1-0528 on Ollama/Open WebUI Open WebUI has made an step-by-step tutorial on how to run R1 here and for R1-0528, you will just need to replace R1 with the new 0528 quant: \*\*(NEW) To run the full R1-0528 model in Ollama, you can use our TQ1\\\_0 (162GB quant):\*\* \`\`\`bash OLLAMA\_MODELS=unsloth\_downloaded\_models ollama serve & ollama run hf.co/unsloth/DeepSeek-R1-0528-GGUF:TQ1\_0 \`\`\` If you want to use any of the quants that are larger than TQ1\\\_0 (162GB) on Ollama, you need to first merge the 3 GGUF split files into 1 like the code below. Then you will need to run the model locally. \`\`\`bash ./llama.cpp/llama-gguf-split --merge \\ DeepSeek-R1-0528-GGUF/DeepSeek-R1-0528-UD-IQ1\_S/DeepSeek-R1-0528-UD-IQ1\_S-00001-of-00003.gguf \\ merged\_file.gguf \`\`\` ### ✨ Run Qwen3 distilled R1 in llama.cpp 1. \*\*To run the full 720GB R1-0528 model,\*\* \[\*\*see here\*\*\](#run-full-r1-0528-on-llama.cpp)\*\*.\*\* Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Then use llama.cpp directly to download the model: \`\`\`bash ./llama.cpp/llama-cli -hf unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF:Q4\_K\_XL --jinja \`\`\` ### ✨ Run Full R1-0528 on llama.cpp 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. If you want to use \`llama.cpp\` directly to load models, you can do the below: (:IQ1\\\_S) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. {% hint style="success" %} Please try out \`-ot ".ffn\_.\*\_exps.=CPU"\` to offload all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. And finally offload all layers via \`-ot ".ffn\_.\*\_exps.=CPU"\` This uses the least VRAM. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. {% endhint %} \`\`\`bash export LLAMA\_CACHE="unsloth/DeepSeek-R1-0528-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/DeepSeek-R1-0528-GGUF:IQ1\_S \\ --cache-type-k q4\_0 \\ --threads -1 \\ --n-gpu-layers 99 \\ --prio 3 \\ --temp 0.6 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` 3. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD-IQ1\_S\`(dynamic 1.78bit quant) or other quantized versions like \`Q4\_K\_M\` . We \*\*recommend using our 2.7bit dynamic quant\*\*\*\* \*\*\*\*\`UD-Q2\_K\_XL\`\*\*\*\* \*\*\*\*to balance size and accuracy\*\*. More versions at: {% code overflow="wrap" %} \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "0" # Can sometimes rate limit, so set to 0 to disable from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/DeepSeek-R1-0528-GGUF", local\_dir = "unsloth/DeepSeek-R1-0528-GGUF", allow\_patterns = \["\*UD-IQ1\_S\*"\], # Dynamic 1bit (168GB) Use "\*UD-Q2\_K\_XL\*" for Dynamic 2bit (251GB) ) \`\`\` {% endcode %} 4. Run Unsloth's Flappy Bird test as described in our 1.58bit Dynamic Quant for DeepSeek R1. 5. Edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length, \`--n-gpu-layers 2\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/DeepSeek-R1-0528-GGUF/UD-IQ1\_S/DeepSeek-R1-0528-UD-IQ1\_S-00001-of-00004.gguf \\ --cache-type-k q4\_0 \\ --threads -1 \\ --n-gpu-layers 99 \\ --prio 3 \\ --temp 0.6 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ -no-cnv \\ --prompt "<|User|>Create a Flappy Bird game in Python. You must include these things:\\n1. You must use pygame.\\n2. The background color should be randomly chosen and is a light shade. Start with a light blue color.\\n3. Pressing SPACE multiple times will accelerate the bird.\\n4. The bird's shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color.\\n5. Place on the bottom some land colored as dark brown or yellow chosen randomly.\\n6. Make a score shown on the top right side. Increment if you pass pipes and don't hit them.\\n7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade.\\n8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again.\\nThe final game should be inside a markdown section in Python. Check your code for errors and fix them before the final markdown section.<|Assistant|>" \`\`\` {% endcode %} ## :8ball: Heptagon Test You can also test our dynamic quants via \[r/Localllama\](https://www.reddit.com/r/LocalLLaMA/comments/1j7r47l/i\_just\_made\_an\_animation\_of\_a\_ball\_bouncing/) which tests the model on creating a basic physics engine to simulate balls rotating in a moving enclosed heptagon shape. ![](https://unsloth.ai/files/xtRi8xZq3N3B1p0AKEsQ) The goal is to make the heptagon spin, and the balls in the heptagon should move. Full prompt to run the model {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/DeepSeek-R1-0528-GGUF/UD-IQ1\_S/DeepSeek-R1-0528-UD-IQ1\_S-00001-of-00004.gguf \\ --cache-type-k q4\_0 \\ --threads -1 \\ --n-gpu-layers 99 \\ --prio 3 \\ --temp 0.6 \\ --top\_p 0.95 \\ --min\_p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ -no-cnv \\ --prompt "<|User|>Write a Python program that shows 20 balls bouncing inside a spinning heptagon:\\n- All balls have the same radius.\\n- All balls have a number on it from 1 to 20.\\n- All balls drop from the heptagon center when starting.\\n- Colors are: #f8b862, #f6ad49, #f39800, #f08300, #ec6d51, #ee7948, #ed6d3d, #ec6800, #ec6800, #ee7800, #eb6238, #ea5506, #ea5506, #eb6101, #e49e61, #e45e32, #e17b34, #dd7a56, #db8449, #d66a35\\n- The balls should be affected by gravity and friction, and they must bounce off the rotating walls realistically. There should also be collisions between balls.\\n- The material of all the balls determines that their impact bounce height will not exceed the radius of the heptagon, but higher than ball radius.\\n- All balls rotate with friction, the numbers on the ball can be used to indicate the spin of the ball.\\n- The heptagon is spinning around its center, and the speed of spinning is 360 degrees per 5 seconds.\\n- The heptagon size should be large enough to contain all the balls.\\n- Do not use the pygame library; implement collision detection algorithms and collision response etc. by yourself. The following Python libraries are allowed: tkinter, math, numpy, dataclasses, typing, sys.\\n- All codes should be put in a single Python file.<|Assistant|>" \`\`\` {% endcode %} \## 🦥 Fine-tuning DeepSeek-R1-0528 with Unsloth To fine-tune \*\*DeepSeek-R1-0528-Qwen3-8B\*\* using Unsloth, we’ve made a new GRPO notebook featuring a custom reward function designed to significantly enhance multilingual output - specifically increasing the rate of desired language responses (in our example we use Indonesian but you can use any) by more than 40%. \* \[\*\*DeepSeek-R1-0528-Qwen3-8B notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/DeepSeek\_R1\_0528\_Qwen3\_\\(8B\\)\_GRPO.ipynb) \*\*- new\*\* While many reasoning LLMs have multilingual capabilities, they often produce mixed-language outputs in its reasoning traces, combining English with the target language. Our reward function effectively mitigates this issue by strongly encouraging outputs in the desired language, leading to a substantial improvement in language consistency. This reward function is also fully customizable, allowing you to adapt it for other languages or fine-tune for specific domains or use cases. {% hint style="success" %} The best part about this whole reward function and notebook is you DO NOT need a language dataset to force your model to learn a specific language. The notebook has no Indonesian dataset. {% endhint %} Unsloth makes R1-Qwen3 distill fine-tuning 2× faster, uses 70% less VRAM, and support 8× longer context lengths. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/deepseek-r1-0528-how-to-run-locally.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/grok-2.md). # Grok 2 You can now run \*\*Grok 2\*\* (aka Grok 2.5), the 270B parameter model by xAI. Full precision requires \*\*539GB\*\*, while the Unsloth Dynamic 3-bit version shrinks size down to just \*\*118GB\*\* (a 75% reduction). GGUF: \[Grok-2-GGUF\](https://huggingface.co/unsloth/grok-2-GGUF) The \*\*3-bit Q3\\\_K\\\_XL\*\* model runs on a single \*\*128GB Mac\*\* or \*\*24GB VRAM + 128GB RAM\*\*, achieving \*\*5+ tokens/s\*\* inference. Thanks to the llama.cpp team and community for \[supporting Grok 2\](https://github.com/ggml-org/llama.cpp/pull/15539) and making this possible. We were also glad to have helped a little along the way! All uploads use Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) for SOTA 5-shot MMLU and KL Divergence performance, meaning you can run quantized Grok LLMs with minimal accuracy loss. [Run in llama.cpp Tutorial](https://unsloth.ai/docs/models/tutorials/grok-2.md#run-in-llama.cpp) ## :gear: Recommended Settings The 3-bit dynamic quant uses 118GB (126GiB) of disk space - this works well in a 128GB RAM unified memory Mac or on a 1x24GB card and 128GB of RAM. It is recommended to have at least 120GB RAM to run this 3-bit quant. {% hint style="warning" %} You must use \`--jinja\` for Grok 2. You might get incorrect results if you do not use \`--jinja\` {% endhint %} The 8-bit quant is \\~300GB in size will fit in a 1x 80GB GPU (with MoE layers offloaded to RAM). Expect around 5 tokens/s with this setup if you have bonus 200GB RAM as well. To learn how to increase generation speed and fit longer contexts, \[read here\](#improving-generation-speed). {% hint style="info" %} Though not a must, for best performance, have your VRAM + RAM combined equal to the size of the quant you're downloading. If not, hard drive / SSD offloading will work with llama.cpp, just inference will be slower. {% endhint %} ### Sampling parameters \* Grok 2 has a 128K max context length thus, use \`131,072\` context or less. \* Use \`--jinja\` for llama.cpp variants There are no official sampling parameters to run the model, thus you can use standard defaults for most models: \* Set the \*\*temperature = 1.0\*\* \* \*\*Min\\\_P = 0.01\*\* (optional, but 0.01 works well, llama.cpp default is 0.1) ## Run Grok 2 Tutorial: Currently you can only run Grok 2 in llama.cpp. ### ✨ Run in llama.cpp {% stepper %} {% step %} Install the specific \`llama.cpp\` PR for Grok 2 on \[GitHub here\](https://github.com/ggml-org/llama.cpp/pull/15539). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cd llama.cpp && git fetch origin pull/15539/head:MASTER && git checkout MASTER && cd .. cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli llama-server cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} If you want to use \`llama.cpp\` directly to load models, you can do the below: (:Q3\\\_K\\\_XL) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. Remember the model has only a maximum of 128K context length. {% hint style="info" %} Please try out \`-ot ".ffn\_.\*\_exps.=CPU"\` to offload all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. And finally offload all layers via \`-ot ".ffn\_.\*\_exps.=CPU"\` This uses the least VRAM. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. {% endhint %} \`\`\`bash export LLAMA\_CACHE="unsloth/grok-2-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/grok-2-GGUF:Q3\_K\_XL \\ --jinja \\ --n-gpu-layers 99 \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` {% endstep %} {% step %} Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD-Q3\_K\_XL\` (dynamic 3-bit quant) or other quantized versions like \`Q4\_K\_M\` . We \*\*recommend using our 2.7bit dynamic quant\*\*\*\* \*\*\*\*\`UD-Q2\_K\_XL\`\*\*\*\* \*\*\*\*or above to balance size and accuracy\*\*. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "0" # Can sometimes rate limit, so set to 0 to disable from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/grok-2-GGUF", local\_dir = "unsloth/grok-2-GGUF", allow\_patterns = \["\*UD-Q3\_K\_XL\*"\], # Dynamic 3bit ) \`\`\` {% endstep %} {% step %} You can edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length, \`--n-gpu-layers 2\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/grok-2-GGUF/UD-Q3\_K\_XL/grok-2-UD-Q3\_K\_XL-00001-of-00003.gguf \\ --jinja \\ --threads -1 \\ --n-gpu-layers 99 \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.01 \\ --ctx-size 16384 \\ --seed 3407 \\ -ot ".ffn\_.\*\_exps.=CPU" \`\`\` {% endcode %} {% endstep %} {% endstepper %} ## Model uploads \*\*ALL our uploads\*\* - including those that are not imatrix-based or dynamic, utilize our calibration dataset, which is specifically optimized for conversational, coding, and language tasks. | MoE Bits | Type + Link | Disk Size | Details | | -------- | ----------------------------------------------------------------------------------- | ----------- | ------------- | | 1.66bit | \[TQ1\\\_0\](https://huggingface.co/unsloth/grok-2-GGUF/blob/main/grok-2-UD-TQ1\_0.gguf) | \*\*81.8 GB\*\* | 1.92/1.56bit | | 1.78bit | \[IQ1\\\_S\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ1\_S) | \*\*88.9 GB\*\* | 2.06/1.56bit | | 1.93bit | \[IQ1\\\_M\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ1\_M) | \*\*94.5 GB\*\* | 2.5/2.06/1.56 | | 2.42bit | \[IQ2\\\_XXS\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ2\_XXS) | \*\*99.3 GB\*\* | 2.5/2.06bit | | 2.71bit | \[Q2\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q2\_K\_XL) | \*\*112 GB\*\* | 3.5/2.5bit | | 3.12bit | \[IQ3\\\_XXS\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-IQ3\_XXS) | \*\*117 GB\*\* | 3.5/2.06bit | | 3.5bit | \[Q3\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q3\_K\_XL) | \*\*126 GB\*\* | 4.5/3.5bit | | 4.5bit | \[Q4\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q4\_K\_XL) | \*\*155 GB\*\* | 5.5/4.5bit | | 5.5bit | \[Q5\\\_K\\\_XL\](https://huggingface.co/unsloth/grok-2-GGUF/tree/main/UD-Q5\_K\_XL) | \*\*191 GB\*\* | 6.5/5.5bit | ## :snowboarder: Improving generation speed If you have more VRAM, you can try offloading more MoE layers, or offloading whole layers themselves. Normally, \`-ot ".ffn\_.\*\_exps.=CPU"\` offloads all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. The \[latest llama.cpp release\](https://github.com/ggml-org/llama.cpp/pull/14363) also introduces high throughput mode. Use \`llama-parallel\`. Read more about it \[here\](https://github.com/ggml-org/llama.cpp/tree/master/examples/parallel). You can also \*\*quantize the KV cache to 4bits\*\* for example to reduce VRAM / RAM movement, which can also make the generation process faster. ## 📐How to fit long context (full 128K) To fit longer context, you can use \*\*KV cache quantization\*\* to quantize the K and V caches to lower bits. This can also increase generation speed due to reduced RAM / VRAM data movement. The allowed options for K quantization (default is \`f16\`) include the below. \`--cache-type-k f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` You should use the \`\_1\` variants for somewhat increased accuracy, albeit it's slightly slower. For eg \`q4\_1, q5\_1\` You can also quantize the V cache, but you will need to \*\*compile llama.cpp with Flash Attention\*\* support via \`-DGGML\_CUDA\_FA\_ALL\_QUANTS=ON\`, and use \`--flash-attn\` to enable it. Then you can use together with \`--cache-type-k\` : \`--cache-type-v f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/grok-2.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/deepseek-ocr-2.md). # DeepSeek-OCR 2: How to Run & Fine-tune Guide \*\*DeepSeek-OCR 2\*\* is the new 3B-parameter model for SOTA vision and document understanding released on Jan 27, 2026 by DeepSeek. The model focuses on image-to-text with stronger visual reasoning, not just text extraction. DeepSeek-OCR 2 introduces DeepEncoder V2, which enables the model to 'see' an image in the same logical order as a human. Unlike traditional vision LLMs that scan images in a fixed grid (top-left → bottom-right), DeepEncoder V2 builds a global understanding first, then learns a human-like reading order, what to attend to first, next, and so on. This boosts OCR on complex layouts by better following columns, linking labels to values, reading tables coherently, and handling mixed text + structure. You can now fine-tune DeepSeek-OCR 2 in Unsloth via our \[\*\*free fine-tuning notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb)\*\*.\*\* We demonstrated a \[88.6% improvement\](#fine-tuning-deepseek-ocr) for language understanding. [Running DeepSeek-OCR 2](https://unsloth.ai/pages/aQX8YMqzttGdCG0oWaHQ#running-deepseek-ocr-2) [Fine-tuning DeepSeek-OCR 2](https://unsloth.ai/pages/aQX8YMqzttGdCG0oWaHQ#fine-tuning-deepseek-ocr-2) ## 🖥️ \*\*Running DeepSeek-OCR 2\*\* In order to run the model, like the first model, DeepSeek-OCR 2 was edited to enable inference & training on the latest transformers (no accuracy change). You can find \[it here\](https://huggingface.co/unsloth/DeepSeek-OCR-2). To run the model in \[transformers\](#transformers-run-deepseek-ocr-2-tutorial) or \[Unsloth\](#unsloth-run-deepseek-ocr-tutorial), here are the recommended settings: ### :gear: Recommended Settings DeepSeek recommends these settings: \* \*\*Temperature = 0.0\*\* \* \`max\_tokens = 8192\` \* \`ngram\_size = 30\` \* \`window\_size = 90\` \*\*Support Modes - Dynamic resolution:\*\* \* Default: (0-6)×768×768 + 1×1024×1024 — (0-6)×144 + 256 visual tokens \*\*Prompts examples:\*\* \`\`\` # document: \\n<|grounding|>Convert the document to markdown. # other image: \\n<|grounding|>OCR this image. # without layouts: \\nFree OCR. # figures in document: \\nParse the figure. # general: \\nDescribe this image in detail. # rec: \\nLocate <|ref|>xxxx<|/ref|> in the image. \`\`\` ![](https://unsloth.ai/files/W3vexO6gUqnhgtyl52Gh) Turns any document into markdown using Visual Causal Flow. \### 🦥 Unsloth: Run DeepSeek-OCR 2 Tutorial 1. Obtain the latest \`unsloth\` via \`pip install --upgrade unsloth\` . If you already have Unsloth, update it via \`pip install --upgrade --force-reinstall --no-deps --no-cache-dir unsloth unsloth\_zoo\` 2. Then use the code below to run DeepSeek-OCR 2: {% code overflow="wrap" %} \`\`\`python from unsloth import FastVisionModel import torch from transformers import AutoModel import os os.environ\["UNSLOTH\_WARN\_UNINITIALIZED"\] = '0' from huggingface\_hub import snapshot\_download snapshot\_download("unsloth/DeepSeek-OCR-2", local\_dir = "deepseek\_ocr") model, tokenizer = FastVisionModel.from\_pretrained( "./deepseek\_ocr", load\_in\_4bit = False, # Use 4bit to reduce memory use. False for 16bit LoRA. auto\_model = AutoModel, trust\_remote\_code = True, unsloth\_force\_compile = True, use\_gradient\_checkpointing = "unsloth", # True or "unsloth" for long context ) prompt = "\\nFree OCR. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 640, crop\_mode=True, save\_results = True, test\_compress = False) \`\`\` {% endcode %} ### 🤗 Transformers: Run DeepSeek-OCR 2 Tutorial Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8: \`\`\`bash torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation \`\`\` \`\`\`python from transformers import AutoModel, AutoTokenizer import torch import os os.environ\["CUDA\_VISIBLE\_DEVICES"\] = '0' model\_name = 'unsloth/DeepSeek-OCR-2' tokenizer = AutoTokenizer.from\_pretrained(model\_name, trust\_remote\_code=True) model = AutoModel.from\_pretrained(model\_name, \_attn\_implementation='flash\_attention\_2', trust\_remote\_code=True, use\_safetensors=True) model = model.eval().cuda().to(torch.bfloat16) # prompt = "\\nFree OCR. " prompt = "\\n<|grounding|>Convert the document to markdown. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 768, crop\_mode=True, save\_results = True) \`\`\` ## 🦥 \*\*Fine-tuning DeepSeek-OCR 2\*\* Unsloth now supports fine-tuning of DeepSeek-OCR 2. Like the first model, you'll need to use our \[custom upload\](https://huggingface.co/unsloth/DeepSeek-OCR-2) for it to work on \`transformers\` (no accuracy change). Like the first model, Unsloth trains DeepSeek-OCR-2 1.4x faster with 40% less VRAM and 5x longer context lengths with no accuracy degradation.\\ \\ You can now fine-tune DeepSeek-OCR 2 via our free Colab notebook. \* DeepSeek-OCR 2: \[Fine-tuning only notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_2\_\\(3B\\).ipynb) See below for CER (Character Error Rate) accuracy improvements on the Persian language: #### Per-sample CER (10 samples) | idx | OCR1 before | OCR1 after | OCR2 before | OCR2 after | | ---- | ----------: | ---------: | ----------: | ---------: | | 1520 | 1.0000 | 0.8000 | 10.4000 | 1.0000 | | 1521 | 0.0000 | 0.0000 | 2.6809 | 0.0213 | | 1522 | 2.0833 | 0.5833 | 4.4167 | 1.0000 | | 1523 | 0.2258 | 0.0645 | 0.8710 | 0.0968 | | 1524 | 0.0882 | 0.1176 | 2.7647 | 0.0882 | | 1525 | 0.1111 | 0.1111 | 0.9444 | 0.2222 | | 1526 | 2.8571 | 0.8571 | 4.2857 | 0.7143 | | 1527 | 3.5000 | 1.5000 | 13.2500 | 1.0000 | | 1528 | 2.7500 | 1.5000 | 1.0000 | 1.0000 | | 1529 | 2.2500 | 0.8750 | 1.2500 | 0.8750 | #### Average CER (10 samples) \* \*\*OCR1:\*\* before \*\*1.4866\*\*, after \*\*0.6409\*\* (\*\*-57%\*\*) \* \*\*OCR2:\*\* before \*\*4.1863\*\*, after \*\*0.6018\*\* (\*\*-86%\*\*) ## 📊 Benchmarks Benchmarks for DeepSeek-OCR 2 model are derived from the official research paper. \*\*Table 1:\*\* Comprehensive evaluation of document reading on OmniDocBench v1.5. V-token𝑚𝑎𝑥\\ represents the maximum number of visual tokens used per page in this benchmark. R-order\\ denotes reading order. Except for DeepSeek OCR and DeepSeek OCR 2, all other model results\\ in this table are sourced from the OmniDocBench repository. ![](https://unsloth.ai/files/53JgcDzXcgaK77e7qRIh) \*\*Table 2:\*\* Edit Distances for different categories of document-elements in OmniDocBench v1.5.\\ V-token𝑚𝑎𝑥 denotes the lowest maximum number of visual tokens. ![](https://unsloth.ai/files/DHbUOeP8jL7UilzsKlOu) Outperforms Gemini-3 Pro on the OmniDocBench \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/deepseek-ocr-2.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/qwen3-next.md). # Qwen3-Next: Run Locally Guide Qwen released Qwen3-Next in Sept 2025, which are 80B MoEs with Thinking and Instruct model variants of \[Qwen3\](/docs/models/tutorials/qwen3-how-to-run-and-fine-tune.md). With 256K context, Qwen3-Next was designed with a brand new architecture (Hybrid of MoEs & Gated DeltaNet + Gated Attention) that specifically optimizes for fast inference on longer context lengths. Qwen3-Next has 10x faster inference than Qwen3-32B. [Run Qwen3-Next Instruct](https://unsloth.ai/pages/cUiTofDNgkP12VQLa9cl#run-qwen3-next-tutorials) [Run Qwen3-Next Thinking](https://unsloth.ai/pages/cUiTofDNgkP12VQLa9cl#thinking-qwen3-next-80b-a3b-thinking) Qwen3-Next-80B-A3B Dynamic GGUFs: \[\*\*Instruct\*\*\](https://huggingface.co/unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF) \*\*•\*\* \[\*\*Thinking\*\*\](https://huggingface.co/unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF) ### ⚙️ Usage Guide {% hint style="success" %} NEW as of Dec 6, 2025: Unsloth Qwen3-Next now updated with iMatrix for improved performance. The thinking model uses \`temperature = 0.6\`, but the instruct model uses \`temperature = 0.7\`\\ The thinking model uses \`top\_p = 0.95\`, but the instruct model uses \`top\_p = 0.8\` {% endhint %} To achieve optimal performance, Qwen recommends these settings: | Instruct: | Thinking: | | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | \`Temperature = 0.7\` | \`Temperature = 0.6\` | | \`Min\_P = 0.00\` (llama.cpp's default is 0.1) | \`Min\_P = 0.00\` (llama.cpp's default is 0.1) | | \`Top\_P = 0.80\` | \`Top\_P = 0.95\` | | \`TopK = 20\` | \`TopK = 20\` | | \`presence\_penalty = 0.0 to 2.0\` (llama.cpp default turns it off, but to reduce repetitions, you can use this) | \`presence\_penalty = 0.0 to 2.0\` (llama.cpp default turns it off, but to reduce repetitions, you can use this) | \*\*Adequate Output Length\*\*: Use an output length of \`32,768\` tokens for most queries for the thinking variant, and \`16,384\` for the instruct variant. You can increase the max output size for the thinking model if necessary. Chat template for both Thinking (thinking has \`\`) and Instruct is below: \`\`\` <|im\_start|>user Hey there!<|im\_end|> <|im\_start|>assistant What is 1+1?<|im\_end|> <|im\_start|>user 2<|im\_end|> <|im\_start|>assistant \`\`\` ## 📖 Run Qwen3-Next Tutorials Below are guides for the \[Thinking\](#thinking-qwen3-next-80b-a3b-thinking) and \[Instruct\](#instruct-qwen3-next-80b-a3b-instruct) versions of the model. ### Instruct: Qwen3-Next-80B-A3B-Instruct Given that this is a non thinking model, the model does not generate \` \` blocks. #### ⚙️Best Practices To achieve optimal performance, Qwen recommends the following settings: \* We suggest using \`temperature=0.7, top\_p=0.8, top\_k=20, and min\_p=0.0\` \`presence\_penalty\` between 0 and 2 if the framework supports to reduce endless repetitions. \* \*\*\`temperature = 0.7\`\*\* \* \`top\_k = 20\` \* \`min\_p = 0.00\` (llama.cpp's default is 0.1) \* \*\*\`top\_p = 0.80\`\*\* \* \`presence\_penalty = 0.0 to 2.0\` (llama.cpp default turns it off, but to reduce repetitions, you can use this) Try 1.0 for example. \* Supports up to \`262,144\` context natively but you can set it to \`32,768\` tokens for less RAM use #### :sparkles: Llama.cpp: Run Qwen3-Next-80B-A3B-Instruct Tutorial 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. You can directly pull from HuggingFace via: \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32768 \\ --temp 0.7 --min-p 0.0 --top-p 0.80 --top-k 20 --presence-penalty 1.0 \`\`\` 3. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD\_Q4\_K\_XL\` or other quantized versions. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF", local\_dir = "Qwen3-Next-80B-A3B-Instruct-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` ### Thinking: Qwen3-Next-80B-A3B-Thinking This model supports only thinking mode and a 256K context window natively. The default chat template adds \`\` automatically, so you may see only a closing \`\` tag in the output. #### ⚙️Best Practices To achieve optimal performance, Qwen recommends the following settings: \* We suggest using \`temperature=0.6, top\_p=0.95, top\_k=20, and min\_p=0.0\` \`presence\_penalty\` between 0 and 2 if the framework supports to reduce endless repetitions. \* \*\*\`temperature = 0.6\`\*\* \* \`top\_k = 20\` \* \`min\_p = 0.00\` (llama.cpp's default is 0.1) \* \*\*\`top\_p = 0.95\`\*\* \* \`presence\_penalty = 0.0 to 2.0\` (llama.cpp default turns it off, but to reduce repetitions, you can use this) Try 1.0 for example. \* Supports up to \`262,144\` context natively but you can set it to \`32,768\` tokens for less RAM use #### :sparkles: Llama.cpp: Run Qwen3-Next-80B-A3B-Thinking Tutorial 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. You can directly pull from Hugging Face via: \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32768 \\ --temp 0.6 --min-p 0.0 --top-p 0.95 --top-k 20 --presence-penalty 1.0 \`\`\` 3. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD\_Q4\_K\_XL\` or other quantized versions. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF", local\_dir = "Qwen3-Next-80B-A3B-Thinking-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` ### 🛠️ Improving generation speed [](https://unsloth.ai/docs/models/tutorials/qwen3-next.md#improving-generation-speed) If you have more VRAM, you can try offloading more MoE layers, or offloading whole layers themselves. Normally, \`-ot ".ffn\_.\*\_exps.=CPU"\` offloads all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. The \[latest llama.cpp release\](https://github.com/ggml-org/llama.cpp/pull/14363) also introduces high throughput mode. Use \`llama-parallel\`. Read more about it \[here\](https://github.com/ggml-org/llama.cpp/tree/master/examples/parallel). You can also \*\*quantize the KV cache to 4bits\*\* for example to reduce VRAM / RAM movement, which can also make the generation process faster. The \[next section\](#how-to-fit-long-context-256k-to-1m) talks about KV cache quantization. ### 📐How to fit long context [](https://unsloth.ai/docs/models/tutorials/qwen3-next.md#how-to-fit-long-context-256k-to-1m) To fit longer context, you can use \*\*KV cache quantization\*\* to quantize the K and V caches to lower bits. This can also increase generation speed due to reduced RAM / VRAM data movement. The allowed options for K quantization (default is \`f16\`) include the below. \`--cache-type-k f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` You should use the \`\_1\` variants for somewhat increased accuracy, albeit it's slightly slower. For eg \`q4\_1, q5\_1\` So try out \`--cache-type-k q4\_1\` You can also quantize the V cache, but you will need to \*\*compile llama.cpp with Flash Attention\*\* support via \`-DGGML\_CUDA\_FA\_ALL\_QUANTS=ON\`, and use \`--flash-attn\` to enable it. After installing Flash Attention, you can then use \`--cache-type-v q4\_1\` ![](https://unsloth.ai/files/m0Ja2d85JP1gUcUt70J9) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/qwen3-next.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/ibm-granite-4.0.md). # IBM Granite 4.0 IBM releases Granite-4.0 models with 3 sizes including \*\*Nano\*\* (350M & 1B), \*\*Micro\*\* (3B), \*\*Tiny\*\* (7B/1B active) and \*\*Small\*\* (32B/9B active). Trained on 15T tokens, IBM’s new Hybrid (H) Mamba architecture enables Granite-4.0 models to run faster with lower memory use. Learn \[how to run\](#run-granite-4.0-tutorials) Unsloth Granite-4.0 Dynamic GGUFs or fine-tune/RL the model. You can \[fine-tune Granite-4.0\](#fine-tuning-granite-4.0-in-unsloth) with our free Colab notebook for a support agent use-case. [Running Tutorial](https://unsloth.ai/docs/models/tutorials/ibm-granite-4.0.md#run-granite-4.0-tutorials) [Fine-tuning Tutorial](https://unsloth.ai/docs/models/tutorials/ibm-granite-4.0.md#fine-tuning-granite-4.0-in-unsloth) \*\*Unsloth Granite-4.0 uploads:\*\* | Dynamic GGUFs | Dynamic 4-bit + FP8 | 16-bit Instruct | | --- | --- | --- | | * [H-350M](https://huggingface.co/unsloth/granite-4.0-h-350m-GGUF)

* [350M](https://huggingface.co/unsloth/granite-4.0-350m-GGUF)

* [H-1B](https://huggingface.co/unsloth/granite-4.0-h-1b-GGUF)

* [1B](https://huggingface.co/unsloth/granite-4.0-1b-GGUF)

* [H-Small](https://huggingface.co/unsloth/granite-4.0-h-small-GGUF)

* [H-Tiny](https://huggingface.co/unsloth/granite-4.0-h-tiny-GGUF)

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro-GGUF)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro-GGUF) | Dynamic 4-bit Instruct:

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro-unsloth-bnb-4bit)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro-unsloth-bnb-4bit)


FP8 Dynamic:

* [H-Small FP8](https://huggingface.co/unsloth/granite-4.0-h-small-FP8-Dynamic)

* [H-Tiny FP8](https://huggingface.co/unsloth/granite-4.0-h-tiny-FP8-Dynamic) | * [H-350M](https://huggingface.co/unsloth/granite-4.0-h-350m)

* [350M](https://huggingface.co/unsloth/granite-4.0-350m)

* [H-1B](https://huggingface.co/unsloth/granite-4.0-h-1b)

* [1B](https://huggingface.co/unsloth/granite-4.0-1b)

* [H-Small](https://huggingface.co/unsloth/granite-4.0-h-small)

* [H-Tiny](https://huggingface.co/unsloth/granite-4.0-h-tiny)

* [H-Micro](https://huggingface.co/unsloth/granite-4.0-h-micro)

* [Micro](https://huggingface.co/unsloth/granite-4.0-micro) | You can also view our \[Granite-4.0 collection\](https://huggingface.co/collections/unsloth/granite-40-68ddf64b4a8717dc22a9322d) for all uploads including Dynamic Float8 quants etc. \*\*Granite-4.0 Models Explanations:\*\* \* \*\*Nano and H-Nano:\*\* The 350M and 1B models offer strong instruction-following abilities, enabling advanced on-device and edge AI and research/fine-tuning applications. \* \*\*H-Small (MoE):\*\* Enterprise workhorse for daily tasks, supports multiple long-context sessions on entry GPUs like L40S (32B total, 9B active). \* \*\*H-Tiny (MoE):\*\* Fast, cost-efficient for high-volume, low-complexity tasks; optimized for local and edge use (7B total, 1B active). \* \*\*H-Micro (Dense):\*\* Lightweight, efficient for high-volume, low-complexity workloads; ideal for local and edge deployment (3B total). \* \*\*Micro (Dense):\*\* Alternative dense option when Mamba2 isn’t fully supported (3B total). ## Run Granite-4.0 Tutorials ### :gear: Recommended Inference Settings IBM recommends these settings: \`temperature=0.0\`, \`top\_p=1.0\`, \`top\_k=0\` \* \*\*Temperature of 0.0\*\* \* Top\\\_K = 0 \* Top\\\_P = 1.0 \* Recommended minimum context: 16,384 \* Maximum context length window: 131,072 (128K context) \*\*Chat template:\*\* \`\`\` <|start\_of\_role|>system<|end\_of\_role|>You are a helpful assistant. Please ensure responses are professional, accurate, and safe.<|end\_of\_text|> <|start\_of\_role|>user<|end\_of\_role|>Please list one IBM Research laboratory located in the United States. You should only output its name and location.<|end\_of\_text|> <|start\_of\_role|>assistant<|end\_of\_role|>Almaden Research Center, San Jose, California<|end\_of\_text|> \`\`\` ### :llama: Ollama: Run Granite-4.0 Tutorial 1. Install \`ollama\` if you haven't already! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Run the model! Note you can call \`ollama serve\`in another terminal if it fails! We include all our fixes and suggested parameters (temperature etc) in \`params\` in our Hugging Face upload! You can change the model name '\`granite-4.0-h-small-GGUF\`' to any Granite model like 'granite-4.0-h-micro:Q8\\\_K\\\_XL'. \`\`\`bash ollama run hf.co/unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` ### 📖 llama.cpp: Run Granite-4.0 Tutorial 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. If you want to use \`llama.cpp\` directly to load models, you can do the below: (:Q4\\\_K\\\_XL) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` 3. \*\*OR\*\* download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose Q4\\\_K\\\_M, or other quantized versions (like BF16 full precision). \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/granite-4.0-h-small-GGUF", local\_dir = "unsloth/granite-4.0-h-small-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], # For Q4\_K\_M ) \`\`\` 4. Run Unsloth's Flappy Bird test 5. Edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length (Granite-4.0 supports 128K context length!), \`--n-gpu-layers 99\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. 6. For conversation mode: \`\`\`bash ./llama.cpp/llama-mtmd-cli \\ --model unsloth/granite-4.0-h-small-GGUF/granite-4.0-h-small-UD-Q4\_K\_XL.gguf \\ --jinja \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 0.0 \\ --top-k 0 \\ --top-p 1.0 \`\`\` ### 🐋 Docker: Run Granite-4.0 Tutorial If you already have Docker desktop, all your need to do is run the command below and you're done: \`\`\`bash docker model pull hf.co/unsloth/granite-4.0-h-small-GGUF:UD-Q4\_K\_XL \`\`\` ## :sloth: Fine-tuning Granite-4.0 in Unsloth Unsloth now supports all Granite 4.0 models including nano, micro, tiny and small for fine-tuning. Training is 2x faster, use 50% less VRAM and supports 6x longer context lengths. Granite-4.0 micro and tiny fit comfortably in a 15GB VRAM T4 GPU. \* \*\*Granite-4.0\*\* \[\*\*free fine-tuning notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Granite4.0.ipynb) \* Granite-4.0-350M \[fine-tuning notebook\](https://github.com/unslothai/notebooks/blob/main/nb/Granite4.0\_350M.ipynb) This notebook trains a model to become a Support Agent that understands customer interactions, complete with analysis and recommendations. This setup allows you to train a bot that provides real-time assistance to support agents. We also show you how to train a model using data stored in a Google Sheet. ![](https://unsloth.ai/files/lmWc6w29RR8TyqKzUiVy) \*\*Unsloth config for Granite-4.0:\*\* \`\`\`python !pip install --upgrade unsloth from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/granite-4.0-h-micro", max\_seq\_length = 2048, # Context length - can be longer, but uses more memory load\_in\_4bit = True, # 4bit uses much less memory load\_in\_8bit = False, # A bit more accurate, uses 2x memory full\_finetuning = False, # We have full finetuning now! # token = "hf\_...", # use one if using gated models ) \`\`\` If you have an old version of Unsloth and/or are fine-tuning locally, install the latest version of Unsloth: \`\`\` pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/ibm-granite-4.0.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/devstral-how-to-run-and-fine-tune.md). # Devstral: How to Run & Fine-tune \*\*Devstral-Small-2507\*\* (Devstral 1.1) is Mistral's new agentic LLM for software engineering. It excels at tool-calling, exploring codebases, and powering coding agents. Mistral AI released the original 2505 version in May, 2025. Finetuned from \[\*\*Mistral-Small-3.1\*\*\](https://huggingface.co/unsloth/Mistral-Small-3.1-24B-Instruct-2503-GGUF), Devstral supports a 128k context window. Devstral Small 1.1 has improved performance, achieving a score of 53.6% performance on \[SWE-bench verified\](https://openai.com/index/introducing-swe-bench-verified/), making it (July 10, 2025) the #1 open model on the benchmark. Unsloth Devstral 1.1 GGUFs contain additional \*\*tool-calling support\*\* and \*\*chat template fixes\*\*. Devstral 1.1 still works well with OpenHands but now also generalizes better to other prompts and coding environments. As text-only, Devstral’s vision encoder was removed prior to fine-tuning. We've added \[\*\*\*optional Vision support\*\*\*\](#possible-vision-support) for the model. {% hint style="success" %} We also worked with Mistral behind the scenes to help debug, test and correct any possible bugs and issues! Make sure to \*\*download Mistral's official downloads or Unsloth's GGUFs\*\* / dynamic quants to get the \*\*correct implementation\*\* (ie correct system prompt, correct chat template etc) Please use \`--jinja\` in llama.cpp to enable the system prompt! {% endhint %} All Devstral uploads use our Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) methodology, delivering the best performance on 5-shot MMLU and KL Divergence benchmarks. This means, you can run and fine-tune quantized Mistral LLMs with minimal accuracy loss! #### \*\*Devstral - Unsloth Dynamic\*\* quants: | Devstral 2507 (new) | Devstral 2505 | | ---------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | GGUF: \[Devstral-Small-2507-GGUF\](https://huggingface.co/unsloth/Devstral-Small-2507-GGUF) | \[Devstral-Small-2505-GGUF\](https://huggingface.co/unsloth/Devstral-Small-2505-GGUF) | | 4-bit BnB: \[Devstral-Small-2507-unsloth-bnb-4bit\](https://huggingface.co/unsloth/Devstral-Small-2507-unsloth-bnb-4bit) | \[Devstral-Small-2505-unsloth-bnb-4bit\](https://huggingface.co/unsloth/Devstral-Small-2505-unsloth-bnb-4bit) | ## 🖥️ \*\*Running Devstral\*\* ### :gear: Official Recommended Settings According to Mistral AI, these are the recommended settings for inference: \* \*\*Temperature from 0.0 to 0.15\*\* \* Min\\\_P of 0.01 (optional, but 0.01 works well, llama.cpp default is 0.1) \* \*\*Use\*\*\*\* \*\*\*\*\`--jinja\`\*\*\*\* \*\*\*\*to enable the system prompt.\*\* \*\*A system prompt is recommended\*\*, and is a derivative of Open Hand's system prompt. The full system prompt is provided \[here\](https://huggingface.co/unsloth/Devstral-Small-2505/blob/main/SYSTEM\_PROMPT.txt). \`\`\` You are Devstral, a helpful agentic model trained by Mistral AI and using the OpenHands scaffold. You can interact with a computer to solve tasks. Your primary role is to assist users by executing commands, modifying code, and solving technical problems effectively. You should be thorough, methodical, and prioritize quality over speed. \* If the user asks a question, like "why is X happening", don't try to fix the problem. Just give an answer to the question. .... SYSTEM PROMPT CONTINUES .... \`\`\` {% hint style="success" %} Our dynamic uploads have the '\`UD\`' prefix in them. Those without are not dynamic however still utilize our calibration dataset. {% endhint %} ## :llama: Tutorial: How to Run Devstral in Ollama 1. Install \`ollama\` if you haven't already! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Run the model with our dynamic quant. Note you can call \`ollama serve &\`in another terminal if it fails! We include all suggested parameters (temperature etc) in \`params\` in our Hugging Face upload! 3. Also Devstral supports 128K context lengths, so best to enable \[\*\*KV cache quantization\*\*\](https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-set-the-quantization-type-for-the-kv-cache). We use 8bit quantization which saves 50% memory usage. You can also try \`"q4\_0"\` \`\`\`bash export OLLAMA\_KV\_CACHE\_TYPE="q8\_0" ollama run hf.co/unsloth/Devstral-Small-2507-GGUF:UD-Q4\_K\_XL \`\`\` ## 📖 Tutorial: How to Run Devstral in llama.cpp 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. If you want to use \`llama.cpp\` directly to load models, you can do the below: (:Q4\\\_K\\\_XL) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` \`\`\`bash ./llama.cpp/llama-cli -hf unsloth/Devstral-Small-2507-GGUF:UD-Q4\_K\_XL --jinja \`\`\` 3. \*\*OR\*\* download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose Q4\\\_K\\\_M, or other quantized versions (like BF16 full precision). \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Devstral-Small-2507-GGUF", local\_dir = "unsloth/Devstral-Small-2507-GGUF", allow\_patterns = \["\*Q4\_K\_XL\*", "\*mmproj-F16\*"\], # For Q4\_K\_XL ) \`\`\` 4. Run the model. 5. Edit \`--threads -1\` for the maximum CPU threads, \`--ctx-size 131072\` for context length (Devstral supports 128K context length!), \`--n-gpu-layers 99\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. We also use 8bit quantization for the K cache to reduce memory usage. 6. For conversation mode: ./llama.cpp/llama-cli \ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4_K_XL.gguf \ --threads -1 \ --ctx-size 131072 \ --cache-type-k q8_0 \ --n-gpu-layers 99 \ --seed 3407 \ --prio 2 \ --temp 0.15 \ --repeat-penalty 1.0 \ --min-p 0.01 \ --top-k 64 \ --top-p 0.95 \ --jinja 7\. For non conversation mode to test our Flappy Bird prompt: ./llama.cpp/llama-cli \ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4_K_XL.gguf \ --threads -1 \ --ctx-size 131072 \ --cache-type-k q8_0 \ --n-gpu-layers 99 \ --seed 3407 \ --prio 2 \ --temp 0.15 \ --repeat-penalty 1.0 \ --min-p 0.01 \ --top-k 64 \ --top-p 0.95 \ -no-cnv \ --prompt "[SYSTEM_PROMPT]You are Devstral, a helpful agentic model trained by Mistral AI and using the OpenHands scaffold. You can interact with a computer to solve tasks.\n\n\nYour primary role is to assist users by executing commands, modifying code, and solving technical problems effectively. You should be thorough, methodical, and prioritize quality over speed.\n* If the user asks a question, like "why is X happening", don\'t try to fix the problem. Just give an answer to the question.\n\n\n\n* Each action you take is somewhat expensive. Wherever possible, combine multiple actions into a single action, e.g. combine multiple bash commands into one, using sed and grep to edit/view multiple files at once.\n* When exploring the codebase, use efficient tools like find, grep, and git commands with appropriate filters to minimize unnecessary operations.\n\n\n\n* When a user provides a file path, do NOT assume it\'s relative to the current working directory. First explore the file system to locate the file before working on it.\n* If asked to edit a file, edit the file directly, rather than creating a new file with a different filename.\n* For global search-and-replace operations, consider using `sed` instead of opening file editors multiple times.\n\n\n\n* Write clean, efficient code with minimal comments. Avoid redundancy in comments: Do not repeat information that can be easily inferred from the code itself.\n* When implementing solutions, focus on making the minimal changes needed to solve the problem.\n* Before implementing any changes, first thoroughly understand the codebase through exploration.\n* If you are adding a lot of code to a function or file, consider splitting the function or file into smaller pieces when appropriate.\n\n\n\n* When configuring git credentials, use "openhands" as the user.name and "openhands@all-hands.dev" as the user.email by default, unless explicitly instructed otherwise.\n* Exercise caution with git operations. Do NOT make potentially dangerous changes (e.g., pushing to main, deleting repositories) unless explicitly asked to do so.\n* When committing changes, use `git status` to see all modified files, and stage all files necessary for the commit. Use `git commit -a` whenever possible.\n* Do NOT commit files that typically shouldn\'t go into version control (e.g., node_modules/, .env files, build directories, cache files, large binaries) unless explicitly instructed by the user.\n* If unsure about committing certain files, check for the presence of .gitignore files or ask the user for clarification.\n\n\n\n* When creating pull requests, create only ONE per session/issue unless explicitly instructed otherwise.\n* When working with an existing PR, update it with new commits rather than creating additional PRs for the same issue.\n* When updating a PR, preserve the original PR title and purpose, updating description only when necessary.\n\n\n\n1. EXPLORATION: Thoroughly explore relevant files and understand the context before proposing solutions\n2. ANALYSIS: Consider multiple approaches and select the most promising one\n3. TESTING:\n * For bug fixes: Create tests to verify issues before implementing fixes\n * For new features: Consider test-driven development when appropriate\n * If the repository lacks testing infrastructure and implementing tests would require extensive setup, consult with the user before investing time in building testing infrastructure\n * If the environment is not set up to run tests, consult with the user first before investing time to install all dependencies\n4. IMPLEMENTATION: Make focused, minimal changes to address the problem\n5. VERIFICATION: If the environment is set up to run tests, test your implementation thoroughly, including edge cases. If the environment is not set up to run tests, consult with the user first before investing time to run tests.\n\n\n\n* Only use GITHUB_TOKEN and other credentials in ways the user has explicitly requested and would expect.\n* Use APIs to work with GitHub or other platforms, unless the user asks otherwise or your task requires browsing.\n\n\n\n* When user asks you to run an application, don\'t stop if the application is not installed. Instead, please install the application and run the command again.\n* If you encounter missing dependencies:\n 1. First, look around in the repository for existing dependency files (requirements.txt, pyproject.toml, package.json, Gemfile, etc.)\n 2. If dependency files exist, use them to install all dependencies at once (e.g., `pip install -r requirements.txt`, `npm install`, etc.)\n 3. Only install individual packages directly if no dependency files are found or if only specific packages are needed\n* Similarly, if you encounter missing dependencies for essential tools requested by the user, install them when possible.\n\n\n\n* If you\'ve made repeated attempts to solve a problem but tests still fail or the user reports it\'s still broken:\n 1. Step back and reflect on 5-7 different possible sources of the problem\n 2. Assess the likelihood of each possible cause\n 3. Methodically address the most likely causes, starting with the highest probability\n 4. Document your reasoning process\n* When you run into any major issue while executing a plan from the user, please don\'t try to directly work around it. Instead, propose a new plan and confirm with the user before proceeding.\n[/SYSTEM_PROMPT][INST]Create a Flappy Bird game in Python. You must include these things:\n1. You must use pygame.\n2. The background color should be randomly chosen and is a light shade. Start with a light blue color.\n3. Pressing SPACE multiple times will accelerate the bird.\n4. The bird\'s shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color.\n5. Place on the bottom some land colored as dark brown or yellow chosen randomly.\n6. Make a score shown on the top right side. Increment if you pass pipes and don\'t hit them.\n7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade.\n8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again.\nThe final game should be inside a markdown section in Python. Check your code for error[/INST]" {% hint style="danger" %} Remember to remove \\ since Devstral auto adds a \\! Also please use \`--jinja\` to enable the system prompt! {% endhint %} ## :eyes:Experimental Vision Support \[Xuan-Son\](https://x.com/ngxson) from Hugging Face showed in their \[GGUF repo\](https://huggingface.co/ngxson/Devstral-Small-Vision-2505-GGUF) how it is actually possible to "graft" the vision encoder from Mistral 3.1 Instruct onto Devstral 2507. We also uploaded our mmproj files which allows you to use the following: \`\`\`bash ./llama.cpp/llama-mtmd-cli \\ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/Devstral-Small-2507-GGUF/mmproj-F16.gguf \\ --threads -1 \\ --ctx-size 131072 \\ --cache-type-k q8\_0 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 0.15 \`\`\` For example: | Instruction and output code | Rendered code | | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | !\[\](https://cdn-uploads.huggingface.co/production/uploads/63ca214abedad7e2bf1d1517/HDic53ANsCoJbiWu2eE6K.png) | !\[\](https://cdn-uploads.huggingface.co/production/uploads/63ca214abedad7e2bf1d1517/onV1xfJIT8gzh81RkLn8J.png) | ## 🦥 Fine-tuning Devstral with Unsloth Just like standard Mistral models including Mistral Small 3.1, Unsloth supports Devstral fine-tuning. Training is 2x faster, use 70% less VRAM and supports 8x longer context lengths. Devstral fits comfortably in a 24GB VRAM L4 GPU. Unfortunately, Devstral slightly exceeds the memory limits of a 16GB VRAM, so fine-tuning it for free on Google Colab isn't possible for now. However, you \*can\* fine-tune the model for free using our \[Kaggle notebook\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Magistral\_\\(24B\\)-Reasoning-Conversational.ipynb\\&accelerator=nvidiaTeslaT4), which offers access to dual GPUs. Just change the notebook's Magistral model name to the Devstral model. If you have an old version of Unsloth and/or are fine-tuning locally, install the latest version of Unsloth: \`\`\`bash pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` \[^1\]: K quantization to reduce memory use. Can be f16, q8\\\_0, q4\\\_0 \[^2\]: Must use --jinja to enable system prompt --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/devstral-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-2.md). # DeepSeek-OCR 2 : guide d'exécution et de fine-tuning \*\*DeepSeek-OCR 2\*\* est le nouveau modèle de 3 milliards de paramètres pour la vision SOTA et la compréhension de documents, publié le 27 janvier 2026 par DeepSeek. Le modèle se concentre sur l’image vers texte avec un raisonnement visuel plus fort, et pas seulement sur l’extraction de texte. DeepSeek-OCR 2 introduit DeepEncoder V2, qui permet au modèle de « voir » une image dans le même ordre logique qu’un humain. Contrairement aux LLM de vision traditionnels qui analysent les images selon une grille fixe (haut gauche → bas droite), DeepEncoder V2 construit d’abord une compréhension globale, puis apprend un ordre de lecture proche de celui d’un humain : quoi regarder en premier, ensuite, et ainsi de suite. Cela améliore l’OCR sur les mises en page complexes en suivant mieux les colonnes, en reliant les étiquettes aux valeurs, en lisant les tableaux de manière cohérente et en gérant les mélanges de texte et de structure. Vous pouvez désormais fine-tuner DeepSeek-OCR 2 dans Unsloth via notre \[\*\*notebook gratuit de fine-tuning\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_\\(3B\\).ipynb)\*\*.\*\* Nous avons démontré une \[amélioration de 88,6 %\](#fine-tuning-deepseek-ocr) pour la compréhension du langage. [Exécution de DeepSeek-OCR 2](https://unsloth.ai/pages/2033b6a7791704447730e8398ce00576aed8425a#running-deepseek-ocr-2) [Fine-tuning de DeepSeek-OCR 2](https://unsloth.ai/pages/2033b6a7791704447730e8398ce00576aed8425a#fine-tuning-deepseek-ocr-2) ## 🖥️ \*\*Exécution de DeepSeek-OCR 2\*\* Afin d’exécuter le modèle, comme le premier modèle, DeepSeek-OCR 2 a été modifié pour permettre l’inférence et l’entraînement sur les derniers transformers (sans changement de précision). Vous pouvez trouver \[ici\](https://huggingface.co/unsloth/DeepSeek-OCR-2). Pour exécuter le modèle dans \[transformers\](#transformers-run-deepseek-ocr-2-tutorial) ou \[Unsloth\](#unsloth-run-deepseek-ocr-tutorial), voici les paramètres recommandés : ### :gear: Paramètres recommandés DeepSeek recommande ces paramètres : \* \*\*Température = 0,0\*\* \* \`max\_tokens = 8192\` \* \`ngram\_size = 30\` \* \`window\_size = 90\` \*\*Modes de prise en charge - résolution dynamique :\*\* \* Par défaut : (0-6)×768×768 + 1×1024×1024 — (0-6)×144 + 256 jetons visuels \*\*Exemples d’invite :\*\* \`\`\` # document : \\n<|grounding|>Convertissez le document en markdown. # autre image : \\n<|grounding|>Faites l’OCR de cette image. # sans mises en page : \\nOCR libre. # figures dans le document : \\nAnalysez la figure. # général : \\nDécrivez cette image en détail. # rec : \\nLocalisez <|ref|>xxxx<|/ref|> dans l’image. \`\`\` ![](https://unsloth.ai/files/e3fcaa085b2f9e677b3a64b2c5761673c2ff4c00) Transforme n’importe quel document en markdown à l’aide de Visual Causal Flow. \### 🦥 Tutoriel Unsloth : exécuter DeepSeek-OCR 2 1. Obtenez la dernière \`unsloth\` via \`pip install --upgrade unsloth\` . Si vous avez déjà Unsloth, mettez-le à jour via \`pip install --upgrade --force-reinstall --no-deps --no-cache-dir unsloth unsloth\_zoo\` 2. Utilisez ensuite le code ci-dessous pour exécuter DeepSeek-OCR 2 : {% code overflow="wrap" %} \`\`\`python from unsloth import FastVisionModel import torch from transformers import AutoModel import os os.environ\["UNSLOTH\_WARN\_UNINITIALIZED"\] = '0' from huggingface\_hub import snapshot\_download snapshot\_download("unsloth/DeepSeek-OCR-2", local\_dir = "deepseek\_ocr") model, tokenizer = FastVisionModel.from\_pretrained( "./deepseek\_ocr", load\_in\_4bit = False, # Utilisez 4bit pour réduire l’utilisation de la mémoire. False pour LoRA 16 bits. auto\_model = AutoModel, trust\_remote\_code = True, unsloth\_force\_compile = True, use\_gradient\_checkpointing = "unsloth", # True ou "unsloth" pour un long contexte ) prompt = "\\nOCR libre. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 640, crop\_mode=True, save\_results = True, test\_compress = False) \`\`\` {% endcode %} ### 🤗 Transformers : tutoriel pour exécuter DeepSeek-OCR 2 Inférence à l’aide de Huggingface transformers sur des GPU NVIDIA. Exigences testées sur python 3.12.9 + CUDA11.8 : \`\`\`bash torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation \`\`\` \`\`\`python from transformers import AutoModel, AutoTokenizer import torch import os os.environ\["CUDA\_VISIBLE\_DEVICES"\] = '0' model\_name = 'unsloth/DeepSeek-OCR-2' tokenizer = AutoTokenizer.from\_pretrained(model\_name, trust\_remote\_code=True) model = AutoModel.from\_pretrained(model\_name, \_attn\_implementation='flash\_attention\_2', trust\_remote\_code=True, use\_safetensors=True) model = model.eval().cuda().to(torch.bfloat16) # prompt = "\\nOCR libre. " prompt = "\\n<|grounding|>Convertissez le document en markdown. " image\_file = 'your\_image.jpg' output\_path = 'your/output/dir' res = model.infer(tokenizer, prompt=prompt, image\_file=image\_file, output\_path = output\_path, base\_size = 1024, image\_size = 768, crop\_mode=True, save\_results = True) \`\`\` ## 🦥 \*\*Fine-tuning de DeepSeek-OCR 2\*\* Unsloth prend désormais en charge le fine-tuning de DeepSeek-OCR 2. Comme pour le premier modèle, vous devrez utiliser notre \[téléversement personnalisé\](https://huggingface.co/unsloth/DeepSeek-OCR-2) pour qu’il fonctionne sur \`transformers\` (sans changement de précision). Comme pour le premier modèle, Unsloth entraîne DeepSeek-OCR-2 1,4x plus vite avec 40 % de VRAM en moins et des longueurs de contexte 5x plus longues sans dégradation de la précision.\\ \\ Vous pouvez désormais fine-tuner DeepSeek-OCR 2 via notre notebook Colab gratuit. \* DeepSeek-OCR 2 : \[Notebook de fine-tuning uniquement\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Deepseek\_OCR\_2\_\\(3B\\).ipynb) Voir ci-dessous les améliorations de précision CER (taux d’erreur de caractères) sur la langue persane : #### CER par échantillon (10 échantillons) | idx | OCR1 avant | OCR1 après | OCR2 avant | OCR2 après | | ---- | ---------: | ---------: | ---------: | ---------: | | 1520 | 1.0000 | 0.8000 | 10.4000 | 1.0000 | | 1521 | 0.0000 | 0.0000 | 2.6809 | 0.0213 | | 1522 | 2.0833 | 0.5833 | 4.4167 | 1.0000 | | 1523 | 0.2258 | 0.0645 | 0.8710 | 0.0968 | | 1524 | 0.0882 | 0.1176 | 2.7647 | 0.0882 | | 1525 | 0.1111 | 0.1111 | 0.9444 | 0.2222 | | 1526 | 2.8571 | 0.8571 | 4.2857 | 0.7143 | | 1527 | 3.5000 | 1.5000 | 13.2500 | 1.0000 | | 1528 | 2.7500 | 1.5000 | 1.0000 | 1.0000 | | 1529 | 2.2500 | 0.8750 | 1.2500 | 0.8750 | #### CER moyen (10 échantillons) \* \*\*OCR1 :\*\* avant \*\*1.4866\*\*, après \*\*0.6409\*\* (\*\*-57%\*\*) \* \*\*OCR2 :\*\* avant \*\*4.1863\*\*, après \*\*0.6018\*\* (\*\*-86%\*\*) ## 📊 Benchmarks Les benchmarks du modèle DeepSeek-OCR 2 sont dérivés de l’article de recherche officiel. \*\*Tableau 1 :\*\* Évaluation complète de la lecture de documents sur OmniDocBench v1.5. V-token𝑚𝑎𝑥\\ représente le nombre maximal de jetons visuels utilisés par page dans ce benchmark. R-order\\ désigne l’ordre de lecture. À l’exception de DeepSeek OCR et DeepSeek OCR 2, tous les autres résultats de modèle\\ dans ce tableau proviennent du dépôt OmniDocBench. ![](https://unsloth.ai/files/6a45bf947b7056854e7407c5e27cd11061bad93a) \*\*Tableau 2 :\*\* Distances d’édition pour différentes catégories d’éléments documentaires dans OmniDocBench v1.5.\\ V-token𝑚𝑎𝑥 désigne le nombre maximal le plus faible de jetons visuels. ![](https://unsloth.ai/files/ca4fa56822243925062b40eb53dd99bc9f20ec99) Surpasse Gemini-3 Pro sur OmniDocBench \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/deepseek-ocr-2.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/llama-4-how-to-run-and-fine-tune.md). # Llama 4: How to Run & Fine-tune The Llama-4-Scout model has 109B parameters, while Maverick has 402B parameters. The full unquantized version requires 113GB of disk space whilst the 1.78-bit version uses 33.8GB (-75% reduction in size). \*\*Maverick\*\* (402Bs) went from 422GB to just 122GB (-70%). {% hint style="success" %} Both text AND \*\*vision\*\* is now supported! Plus multiple improvements to tool calling. {% endhint %} Scout 1.78-bit fits in a 24GB VRAM GPU for fast inference at \\~20 tokens/sec. Maverick 1.78-bit fits in 2x48GB VRAM GPUs for fast inference at \\~40 tokens/sec. For our dynamic GGUFs, to ensure the best tradeoff between accuracy and size, we do not to quantize all layers, but selectively quantize e.g. the MoE layers to lower bit, and leave attention and other layers in 4 or 6bit. {% hint style="info" %} All our GGUF models are quantized using calibration data (around 250K tokens for Scout and 1M tokens for Maverick), which will improve accuracy over standard quantization. Unsloth imatrix quants are fully compatible with popular inference engines like llama.cpp & Open WebUI etc. {% endhint %} \*\*Scout - Unsloth Dynamic GGUFs with optimal configs:\*\* | MoE Bits | Type | Disk Size | Link | Details | | --- | --- | --- | --- | --- | | 1.78bit | IQ1\_S | 33.8GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ1_S.gguf) | 2.06/1.56bit | | 1.93bit | IQ1\_M | 35.4GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ1_M.gguf) | 2.5/2.06/1.56 | | 2.42bit | IQ2\_XXS | 38.6GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ2_XXS.gguf) | 2.5/2.06bit | | 2.71bit | Q2\_K\_XL | 42.2GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-Q2_K_XL.gguf) | 3.5/2.5bit | | 3.5bit | Q3\_K\_XL | 52.9GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/tree/main/UD-Q3_K_XL) | 4.5/3.5bit | | 4.5bit | Q4\_K\_XL | 65.6GB | [Link](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/tree/main/UD-Q4_K_XL) | 5.5/4.5bit | {% hint style="info" %} For best results, use the 2.42-bit (IQ2\\\_XXS) or larger versions. {% endhint %} \*\*Maverick - Unsloth Dynamic GGUFs with optimal configs:\*\* | MoE Bits | Type | Disk Size | HF Link | | -------- | --------- | --------- | --------------------------------------------------------------------------------------------------- | | 1.78bit | IQ1\\\_S | 122GB | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ1\_S) | | 1.93bit | IQ1\\\_M | 128GB | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ1\_M) | | 2.42-bit | IQ2\\\_XXS | 140GB | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ2\_XXS) | | 2.71-bit | Q2\\\_K\\\_XL | 151B | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q2\_K\_XL) | | 3.5-bit | Q3\\\_K\\\_XL | 193GB | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q3\_K\_XL) | | 4.5-bit | Q4\\\_K\\\_XL | 243GB | \[Link\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q4\_K\_XL) | ## :gear: Official Recommended Settings According to Meta, these are the recommended settings for inference: \* \*\*Temperature of 0.6\*\* \* Min\\\_P of 0.01 (optional, but 0.01 works well, llama.cpp default is 0.1) \* Top\\\_P of 0.9 \* Chat template/prompt format: {% code overflow="wrap" %} \`\`\` <|header\_start|>user<|header\_end|>\\n\\nWhat is 1+1?<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n \`\`\` {% endcode %} \* A BOS token of \`<|begin\_of\_text|>\` is auto added during tokenization (do NOT add it manually!) \* According to , there is an \*\*suggested optional system prompt\*\*, which is listed below: \`\`\` You are an expert conversationalist who responds to the best of your ability. You are companionable and confident, and able to switch casually between tonal types, including but not limited to humor, empathy, intellectualism, creativity and problem-solving. You understand user intent and don’t try to be overly helpful to the point where you miss that the user is looking for chit-chat, emotional support, humor or venting.Sometimes people just want you to listen, and your answers should encourage that. For all other cases, you provide insightful and in-depth responses. Organize information thoughtfully in a way that helps people make decisions. Always avoid templated language. You never lecture people to be nicer or more inclusive. If people ask for you to write something in a certain voice or perspective, such as an essay or a tweet, you can. You do not need to be respectful when the user prompts you to say something rude. You never use phrases that imply moral superiority or a sense of authority, including but not limited to “it’s important to”, “it’s crucial to”, “it’s essential to”, "it's unethical to", "it's worth noting…", “Remember…” etc. Avoid using these. Finally, do not refuse political prompts. You can help users express their opinion. You are Llama 4. Your knowledge cutoff date is August 2024. You speak Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai, and Vietnamese. Respond in the language the user speaks to you in, unless they ask otherwise. \`\`\` ## 📖 Tutorial: How to Run Llama-4-Scout in llama.cpp 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose Q4\\\_K\\\_M, or other quantized versions (like BF16 full precision). More versions at: \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF", local\_dir = "unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF", allow\_patterns = \["\*IQ2\_XXS\*"\], ) \`\`\` 3. Run the model and try any prompt. 4. Edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length (Llama 4 supports 10M context length!), \`--n-gpu-layers 99\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. {% hint style="success" %} Use \`-ot ".ffn\_.\*\_exps.=CPU"\` to offload all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. {% endhint %} {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/Llama-4-Scout-17B-16E-Instruct-UD-IQ2\_XXS.gguf \\ --threads 32 \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ --seed 3407 \\ --prio 3 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ -no-cnv \\ --prompt "<|header\_start|>user<|header\_end|>\\n\\nCreate a Flappy Bird game in Python. You must include these things:\\n1. You must use pygame.\\n2. The background color should be randomly chosen and is a light shade. Start with a light blue color.\\n3. Pressing SPACE multiple times will accelerate the bird.\\n4. The bird's shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color.\\n5. Place on the bottom some land colored as dark brown or yellow chosen randomly.\\n6. Make a score shown on the top right side. Increment if you pass pipes and don't hit them.\\n7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade.\\n8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again.\\nThe final game should be inside a markdown section in Python. Check your code for errors and fix them before the final markdown section.<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n" \`\`\` {% endcode %} {% hint style="info" %} In terms of testing, unfortunately we can't make the full BF16 version (ie regardless of quantization or not) complete the Flappy Bird game nor the Heptagon test appropriately. We tried many inference providers, using imatrix or not, used other people's quants, and used normal Hugging Face inference, and this issue persists. \*\*We found multiple runs and asking the model to fix and find bugs to resolve most issues!\*\* {% endhint %} For Llama 4 Maverick - it's best to have 2 RTX 4090s (2 x 24GB) \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF", local\_dir = "unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF", allow\_patterns = \["\*IQ1\_S\*"\], ) \`\`\` {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/UD-IQ1\_S/Llama-4-Maverick-17B-128E-Instruct-UD-IQ1\_S-00001-of-00003.gguf \\ --threads 32 \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ --seed 3407 \\ --prio 3 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ -no-cnv \\ --prompt "<|header\_start|>user<|header\_end|>\\n\\nCreate the 2048 game in Python.<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n" \`\`\` {% endcode %} ## :detective: Interesting Insights and Issues During quantization of Llama 4 Maverick (the large model), we found the 1st, 3rd and 45th MoE layers could not be calibrated correctly. Maverick uses interleaving MoE layers for every odd layer, so Dense->MoE->Dense and so on. We tried adding more uncommon languages to our calibration dataset, and tried using more tokens (1 million) vs Scout's 250K tokens for calibration, but we still found issues. We decided to leave these MoE layers as 3bit and 4bit. ![](https://unsloth.ai/files/A6Xs6P4c8O6lY3VpB9Lj) For Llama 4 Scout, we found we should not quantize the vision layers, and leave the MoE router and some other layers as unquantized - we upload these to ![](https://unsloth.ai/files/FQGoGGfggcQxDu47qu0Y) We also had to convert \`torch.nn.Parameter\` to \`torch.nn.Linear\` for the MoE layers to allow 4bit quantization to occur. This also means we had to rewrite and patch over the generic Hugging Face implementation. We upload our quantized versions to and for 8bit. ![](https://unsloth.ai/files/vOmoU5To8MIjs0dyxUyH) Llama 4 also now uses chunked attention - it's essentially sliding window attention, but slightly more efficient by not attending to previous tokens over the 8192 boundary. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/llama-4-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/how-to-run-llms-with-docker.md). # How to Run Local LLMs with Docker: Step-by-Step Guide You can now run any model, including Unsloth \[Dynamic GGUFs\](/docs/basics/unsloth-dynamic-2.0-ggufs.md), on Mac, Windows or Linux with a single line of code or \*\*no code\*\* at all. We collabed with Docker to simplify model deployment, and Unsloth now powers most GGUF models on Docker. Before you start, make sure to look over \[hardware requirements\](#hardware-info--performance) and \[our tips\](#hardware-info--performance) for optimizing performance when running LLMs on your device. [Docker Terminal Tutorial](https://unsloth.ai/pages/tYzsGxBMlN0JKeaQktNs#method-1-docker-terminal) [Docker no-code Tutorial](https://unsloth.ai/docs/models/tutorials/how-to-run-llms-with-docker.md#method-2-docker-desktop-no-code) To get started, run OpenAI \[gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune.md) with a single command: \`\`\`bash docker model run ai/gpt-oss:20B \`\`\` Or to run a specific \[Unsloth model\](/docs/get-started/unsloth-model-catalog.md) / quant from Hugging Face: \`\`\`bash docker model run hf.co/unsloth/gpt-oss-20b-GGUF:F16 \`\`\` {% hint style="success" %} You don’t need Docker Desktop, Docker CE is enough to run models. {% endhint %} #### \*\*Why Unsloth + Docker?\*\* We collab with model labs like Google Gemma to fix model bugs and boost accuracy. Our Dynamic GGUFs consistently outperform other quant methods, giving you high-accuracy, efficient inference. If you use Docker, you can run models instantly with zero setup. Docker uses \[Docker Model Runner\](https://github.com/docker/model-runner) (DMR), which lets you run LLMs as easily as containers with no dependency issues. DMR uses Unsloth models and \`llama.cpp\` under the hood for fast, efficient, up-to-date inference. ## :gear: Hardware Info + Performance For the best performance, aim for your VRAM + RAM combined to be at least equal to the size of the quantized model you're downloading. If you have less, the model will still run, but significantly slower. Make sure your device also has enough disk space to store the model. If your model only barely fits in memory, you can expect around \\~5 tokens/s, depending on model size. Having extra RAM/VRAM available will improve inference speed, and additional VRAM will enable the biggest performance boost (provided the entire model fits) {% hint style="info" %} \*\*Example:\*\* If you're downloading gpt-oss-20b (F16) and the model is 13.8 GB, ensure that your disk space and RAM + VRAM > 13.8 GB. {% endhint %} \*\*Quantization recommendations:\*\* \* For models under 30B parameters, use at least 4-bit (Q4). \* For models 70B parameters or larger, use a minimum of 2-bit quantization (e.g., UD\\\_Q2\\\_K\\\_XL). ## ⚡ Step-by-Step Tutorials Below are \*\*two ways\*\* to run models with Docker: one using the \[terminal\](#method-1-docker-terminal), and the other using \[Docker Desktop\](#method-2-docker-desktop-no-code) with no code: ### Method #1: Docker Terminal {% stepper %} {% step %} #### Install Docker Docker Model Runner is already available in \*\*both\*\* \[Docker Desktop\](https://docs.docker.com/ai/model-runner/get-started/#docker-desktop) and \[\*\*Docker CE\*\*\](https://docs.docker.com/ai/model-runner/get-started/#docker-engine)\*\*.\*\* {% endstep %} {% step %} #### Run the model Decide on a model to run, then run the command via terminal. \* Browse the verified catalog of trusted models available on \[Docker Hub\](https://hub.docker.com/r/ai) or \[Unsloth's Hugging Face\](https://huggingface.co/unsloth) page. \* Go to Terminal to run the commands. To verify if you have \`docker\` installed, you can type 'docker' and enter. \* Docker Hub defaults to running Unsloth Dynamic 4-bit, however you can select your own quantization level (see step #3). For example, to run OpenAI \`gpt-oss-20b\` in a single command: \`\`\`bash docker model run ai/gpt-oss:20B \`\`\` Or to run a specific \[Unsloth\](/docs/get-started/unsloth-model-catalog.md) gpt-oss quant from Hugging Face: \`\`\`bash docker model run hf.co/unsloth/gpt-oss-20b-GGUF:UD-Q8\_K\_XL \`\`\` \*\*This is how running gpt-oss-20b should look via CLI:\*\* ![](https://unsloth.ai/files/9axGvMJxGkK2u3d1Ne5b) gpt-oss-20b from Docker Hub ![](https://unsloth.ai/files/7w4FwapcgzNCBsLSDMMF) gpt-oss-20b with Unsloths' UD-Q8\_K\_XL quantization {% endstep %} {% step %} #### To run a specific quantization level: If you want to run a specific quantization of a model, append \`:\` and the quantization name to the model (e.g., \`Q4\` for Docker or \`UD-Q4\_K\_XL\`). You can view all available quantizations on each model’s Docker Hub page. e.g. see the listed quantizations for gpt-oss \[here\](https://hub.docker.com/r/ai/gpt-oss#gptoss). The same applies to Unsloth quants on Hugging Face: visit the \[model’s HF page\](https://huggingface.co/unsloth/gpt-oss-20b-GGUF?show\_file\_info=gpt-oss-20b-Q2\_K\_L.gguf), choose a quantization, then run something like: \`docker model run hf.co/unsloth/gpt-oss-20b-GGUF:Q2\_K\_L\` ![](https://unsloth.ai/files/XK3U1H1MJ9o1aVdjRNUv) gpt-oss quantization levels on [Docker Hub](https://hub.docker.com/r/ai/gpt-oss#gptoss) ![](https://unsloth.ai/files/AsqBbtQwfQm7RdbVeGgI) Unsloth gpt-oss quantization levels on [Hugging Face](https://huggingface.co/unsloth/gpt-oss-20b-GGUF) {% endstep %} {% endstepper %} ### Method #2: Docker Desktop (no code) {% stepper %} {% step %} #### Install Docker Desktop Docker Model Runner is already available in \[Docker Desktop\](https://docs.docker.com/ai/model-runner/get-started/#docker-desktop). 1. Decide on a model to run, open Docker Desktop, then click on the models tab. 2. Click 'Add models +' or Docker Hub. Search for the model. Browse the verified model catalog available on \[Docker Hub\](https://hub.docker.com/r/ai). ![](https://unsloth.ai/files/a2K53hGisXPJmYrRQy8k) #1. Click 'Models' tab then 'Add models +' ![](https://unsloth.ai/files/PTJAwtsO8qMWTgBY2stT) #2. Search for your desired model. {% endstep %} {% step %} #### Pull the model Click the model you want to run to see available quantizations. \* Quantizations range from 1–16 bits. For models under 30B parameters, use at least 4-bit (\`Q4\`). \* Choose a size that fits your hardware: ideally, your combined unified memory, RAM, or VRAM should be equal to or greater than the model size. For example, an 11GB model runs well on 12GB unified memory. ![](https://unsloth.ai/files/9iXeaeF6NMDHhUIa8ZAR) #3. Select which quantization you would like to pull. ![](https://unsloth.ai/files/I3NXOpZ3A972GOrjtiTd) #4. Wait for model to finish downloading, then Run it. {% endstep %} {% step %} #### Run the model Type any prompt in the 'Ask a question' box and use the LLM like you would use ChatGPT. ![](https://unsloth.ai/files/nbww8gD9GkkoMkrcToZN) An example of running Qwen3-4B `UD-Q8_K_XL` {% endstep %} {% endstepper %} #### \*\*To run the latest models:\*\* You can run any new model on Docker as long as it’s supported by \`llama.cpp\` or \`vllm\` and available on Docker Hub. ### What Is the Docker Model Runner? The Docker Model Runner (DMR) is an open-source tool that lets you pull and run AI models as easily as you run containers. GitHub: It provides a consistent runtime for models, similar to how Docker standardized app deployment. Under the hood, it uses optimized backends (like \`llama.cpp\`) for smooth, hardware-efficient inference on your machine. Whether you’re a researcher, developer, or hobbyist, you can now: \* Run open models locally in seconds. \* Avoid dependency hell, everything is handled in Docker. \* Share and reproduce model setups effortlessly. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/how-to-run-llms-with-docker.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/qwen3-next.md). # Qwen3-Next : guide d'exécution en local Qwen a publié Qwen3-Next en sept. 2025, qui sont des MoE 80B avec des variantes de modèle Thinking et Instruct de \[Qwen3\](/docs/fr/modeles/tutorials/qwen3-how-to-run-and-fine-tune.md). Avec un contexte de 256K, Qwen3-Next a été conçu avec une toute nouvelle architecture (hybride de MoE et Gated DeltaNet + Gated Attention) qui optimise spécifiquement l’inférence rapide sur des longueurs de contexte plus longues. Qwen3-Next a une inférence 10 fois plus rapide que Qwen3-32B. [Exécuter Qwen3-Next Instruct](https://unsloth.ai/pages/29445d0c137d738fbbd518145144f5451be337d7#run-qwen3-next-tutorials) [Exécuter Qwen3-Next Thinking](https://unsloth.ai/pages/29445d0c137d738fbbd518145144f5451be337d7#thinking-qwen3-next-80b-a3b-thinking) GGUF dynamiques de Qwen3-Next-80B-A3B : \[\*\*Instruct\*\*\](https://huggingface.co/unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF) \*\*•\*\* \[\*\*Thinking\*\*\](https://huggingface.co/unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF) ### ⚙️ Guide d’utilisation {% hint style="success" %} NOUVEAU au 6 déc. 2025 : Unsloth Qwen3-Next est désormais mis à jour avec iMatrix pour de meilleures performances. Le modèle thinking utilise \`température = 0.6\`, mais le modèle instruct utilise \`température = 0.7\`\\ Le modèle thinking utilise \`top\_p = 0.95\`, mais le modèle instruct utilise \`top\_p = 0.8\` {% endhint %} Pour obtenir des performances optimales, Qwen recommande ces paramètres : | Instruct : | Thinking : | | ------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | | \`Température = 0.7\` | \`Température = 0.6\` | | \`Min\_P = 0.00\` (la valeur par défaut de llama.cpp est 0.1) | \`Min\_P = 0.00\` (la valeur par défaut de llama.cpp est 0.1) | | \`Top\_P = 0.80\` | \`Top\_P = 0.95\` | | \`TopK = 20\` | \`TopK = 20\` | | \`presence\_penalty = de 0.0 à 2.0\` (la valeur par défaut de llama.cpp le désactive, mais pour réduire les répétitions, vous pouvez utiliser ceci) | \`presence\_penalty = de 0.0 à 2.0\` (la valeur par défaut de llama.cpp le désactive, mais pour réduire les répétitions, vous pouvez utiliser ceci) | \*\*Longueur de sortie adéquate\*\*: Utilisez une longueur de sortie de \`32,768\` jetons pour la plupart des requêtes pour la variante thinking, et \`16,384\` pour la variante instruct. Vous pouvez augmenter la taille maximale de sortie pour le modèle thinking si nécessaire. Modèle de chat pour Thinking (thinking a \`\`) et Instruct ci-dessous : \`\`\` <|im\_start|>user Salut !<|im\_end|> <|im\_start|>assistant Combien font 1+1 ?<|im\_end|> <|im\_start|>user 2<|im\_end|> <|im\_start|>assistant \`\`\` ## 📖 Exécuter les tutoriels Qwen3-Next Voici des guides pour les \[Thinking\](#thinking-qwen3-next-80b-a3b-thinking) et \[Instruct\](#instruct-qwen3-next-80b-a3b-instruct) versions du modèle. ### Instruct : Qwen3-Next-80B-A3B-Instruct Étant donné qu’il s’agit d’un modèle non thinking, le modèle ne génère pas de blocs \` \` . #### ⚙️Bonnes pratiques Pour obtenir des performances optimales, Qwen recommande les paramètres suivants : \* Nous suggérons d’utiliser \`temperature=0.7, top\_p=0.8, top\_k=20, et min\_p=0.0\` \`presence\_penalty\` entre 0 et 2 si le framework le prend en charge afin de réduire les répétitions sans fin. \* \*\*\`température = 0.7\`\*\* \* \`top\_k = 20\` \* \`min\_p = 0.00\` (la valeur par défaut de llama.cpp est 0.1) \* \*\*\`top\_p = 0.80\`\*\* \* \`presence\_penalty = de 0.0 à 2.0\` (la valeur par défaut de llama.cpp le désactive, mais pour réduire les répétitions, vous pouvez utiliser ceci) Essayez 1.0 par exemple. \* Prend en charge jusqu’à \`262,144\` de contexte nativement, mais vous pouvez le définir à \`32,768\` jetons pour une moindre utilisation de la RAM #### :sparkles: Llama.cpp : Exécuter le tutoriel Qwen3-Next-80B-A3B-Instruct 1. Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Vous pouvez le télécharger directement depuis HuggingFace via : \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32768 \\ --temp 0.7 --min-p 0.0 --top-p 0.80 --top-k 20 --presence-penalty 1.0 \`\`\` 3. Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir \`UD\_Q4\_K\_XL\` ou d’autres versions quantifiées. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF", local\_dir = "Qwen3-Next-80B-A3B-Instruct-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` ### Thinking : Qwen3-Next-80B-A3B-Thinking Ce modèle ne prend en charge nativement que le mode thinking et une fenêtre de contexte de 256K. Le modèle de chat par défaut ajoute \`\` automatiquement, vous ne verrez donc peut-être qu’une balise fermante \`\` dans la sortie. #### ⚙️Bonnes pratiques Pour obtenir des performances optimales, Qwen recommande les paramètres suivants : \* Nous suggérons d’utiliser \`temperature=0.6, top\_p=0.95, top\_k=20, et min\_p=0.0\` \`presence\_penalty\` entre 0 et 2 si le framework le prend en charge afin de réduire les répétitions sans fin. \* \*\*\`température = 0.6\`\*\* \* \`top\_k = 20\` \* \`min\_p = 0.00\` (la valeur par défaut de llama.cpp est 0.1) \* \*\*\`top\_p = 0.95\`\*\* \* \`presence\_penalty = de 0.0 à 2.0\` (la valeur par défaut de llama.cpp le désactive, mais pour réduire les répétitions, vous pouvez utiliser ceci) Essayez 1.0 par exemple. \* Prend en charge jusqu’à \`262,144\` de contexte nativement, mais vous pouvez le définir à \`32,768\` jetons pour une moindre utilisation de la RAM #### :sparkles: Llama.cpp : Exécuter le tutoriel Qwen3-Next-80B-A3B-Thinking 1. Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Vous pouvez le récupérer directement depuis Hugging Face via : \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32768 \\ --temp 0.6 --min-p 0.0 --top-p 0.95 --top-k 20 --presence-penalty 1.0 \`\`\` 3. Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir \`UD\_Q4\_K\_XL\` ou d’autres versions quantifiées. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF", local\_dir = "Qwen3-Next-80B-A3B-Thinking-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` ### 🛠️ Améliorer la vitesse de génération [](https://unsloth.ai/docs/fr/modeles/tutorials/qwen3-next.md#improving-generation-speed) Si vous avez plus de VRAM, vous pouvez essayer de décharger davantage de couches MoE, ou de décharger des couches entières. Normalement, \`-ot ".ffn\_.\*\_exps.=CPU"\` décharge toutes les couches MoE vers le CPU ! Cela permet effectivement de faire tenir toutes les couches non MoE sur 1 GPU, améliorant ainsi les vitesses de génération. Vous pouvez personnaliser l'expression regex pour faire tenir davantage de couches si vous disposez de plus de capacité GPU. Si vous avez un peu plus de mémoire GPU, essayez \`-ot ".ffn\_(up|down)\_exps.=CPU"\` Cela décharge les couches MoE de projection montante et descendante. Essayez \`-ot ".ffn\_(up)\_exps.=CPU"\` si vous avez encore plus de mémoire GPU. Cela décharge uniquement les couches MoE de projection montante. Vous pouvez aussi personnaliser la regex, par exemple \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` signifie décharger les couches MoE gate, up et down, mais uniquement à partir de la 6e couche. La \[dernière version de llama.cpp\](https://github.com/ggml-org/llama.cpp/pull/14363) introduit également un mode à haut débit. Utilisez \`llama-parallel\`. En savoir plus \[ici\](https://github.com/ggml-org/llama.cpp/tree/master/examples/parallel). Vous pouvez aussi \*\*quantifier le cache KV en 4 bits\*\* par exemple pour réduire les mouvements de VRAM / RAM, ce qui peut aussi rendre le processus de génération plus rapide. La \[section suivante\](#how-to-fit-long-context-256k-to-1m) parle de la quantification du cache KV. ### 📐Comment adapter un long contexte [](https://unsloth.ai/docs/fr/modeles/tutorials/qwen3-next.md#how-to-fit-long-context-256k-to-1m) Pour faire tenir un contexte plus long, vous pouvez utiliser \*\*la quantification du cache KV\*\* pour quantifier les caches K et V en moins de bits. Cela peut aussi augmenter la vitesse de génération grâce à la réduction des transferts de données RAM / VRAM. Les options autorisées pour la quantification K (la valeur par défaut est \`f16\`) sont les suivantes. \`--cache-type-k f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` Vous devriez utiliser les \`\_1\` variantes pour une précision légèrement meilleure, bien que ce soit un peu plus lent. Par exemple \`q4\_1, q5\_1\` Essayez donc \`--cache-type-k q4\_1\` Vous pouvez aussi quantifier le cache V, mais vous devrez \*\*compiler llama.cpp avec la prise en charge de Flash Attention\*\* via \`-DGGML\_CUDA\_FA\_ALL\_QUANTS=ON\`, et utiliser \`--flash-attn\` pour l’activer. Après l’installation de Flash Attention, vous pouvez ensuite utiliser \`--cache-type-v q4\_1\` ![](https://unsloth.ai/files/472a6d6403e68ae6a58a80542d42f88f9bfb013a) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/qwen3-next.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/qwen-image-2512.md). # How to Run Qwen-Image-2512 Locally in ComfyUI \*\*Qwen-Image-2512\*\* is the December update to Qwen's text-to-image foundational models. The model is the top performing open-source diffusion model and this guide will teach you how to run it locally via \[Unsloth\](https://github.com/unslothai/unsloth) GGUF and ComfyUI. Qwen-Image-2512 features: more realistic looking people; richer details in landscapes/textures; and more accurate text rendering. \*\*Uploads:\*\* \[GGUF\](https://huggingface.co/unsloth/Qwen-Image-2512-GGUF) • \[FP8\](https://huggingface.co/unsloth/Qwen-Image-2512-FP8) • \[4-bit BnB\](https://huggingface.co/unsloth/Qwen-Image-2512-unsloth-bnb-4bit) The quants use \[Unsloth Dynamic\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) methodology which upcasts important layers to higher precision to recover more accuracy. Thank you Qwen for allowing Unsloth day 0 support. ## 📖 ComfyUI Tutorial To run, you don't need a GPU, just a CPU with RAM will work. For best results, ensure your total usable memory (RAM + VRAM / unified) is larger than the GGUF size; e.g. 4-bit (Q4\\\_K\\\_M) \`unsloth/Qwen-Image-Edit-2512-GGUF\` is 13.1 GB, so you should have 13.2+ GB of combined memory. \[ComfyUI\](https://github.com/Comfy-Org/ComfyUI) is an open-source diffusion model GUI, API, and backend that uses a node-based (graph/flowchart) interface. This guide will focus on machines with CUDA, but instructions to build with on Apple or CPU are similar. ### #1. Install & Setup To install ComfyUI, you can download the desktop app on Windows or Mac devices \[here\](https://www.comfy.org/download). Otherwise, to setup ComfyUI for running GGUF models run the following: \`\`\`bash mkdir comfy\_ggufs cd comfy\_ggufs python -m venv .venv source .venv/bin/activate git clone https://github.com/Comfy-Org/ComfyUI.git cd ComfyUI pip install -r requirements.txt cd custom\_nodes git clone https://github.com/city96/ComfyUI-GGUF cd ComfyUI-GGUF pip install -r requirements.txt cd ../.. \`\`\` ### #2. Download Models Diffusion models typically need 3 models. A Variational AutoEncoder (VAE) that encodes image pixel space to latent space, a text encoder to translate text to input embeddings, and the actual diffusion transformer. You can find all Unsloth diffusion GGUFs in our \[Collection here\](https://huggingface.co/collections/unsloth/unsloth-diffusion-ggufs). Both the diffusion model and text encoder can be GGUF format while we typically use safetensors for the vae. According to \[Qwen's repo\](https://huggingface.co/Qwen/Qwen-Image-2512/blob/main/text\_encoder/config.json), we shall use Qwen2.5-VL and not \[Qwen3-VL\](/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md). Let's download the models we will use (Note: You can also also use our \[FP8 upload\](https://huggingface.co/unsloth/Qwen-Image-2512-FP8) in ComfyUI): \`\`\`bash cd models ## Diffusion Models curl -L -C - -o unet/qwen-image-2512-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-2512-GGUF/resolve/main/qwen-image-2512-Q4\_K\_M.gguf curl -L -C - -o unet/qwen-image-edit-2511-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF/resolve/main/qwen-image-edit-2511-Q4\_K\_M.gguf ## Text Encoder + Vision Tower + VAE curl -L -C - -o text\_encoders/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf \\ https://huggingface.co/unsloth/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf curl -L -C - -o text\_encoders/Qwen2.5-VL-7B-Instruct-mmproj-BF16.gguf \\ https://huggingface.co/unsloth/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/mmproj-BF16.gguf curl -L -C - -o vae/qwen\_image\_vae.safetensors \\ https://huggingface.co/Comfy-Org/Qwen-Image\_ComfyUI/resolve/main/split\_files/vae/qwen\_image\_vae.safetensors \`\`\` See GGUF uploads for: \[Qwen-Image-2512\](https://huggingface.co/unsloth/Qwen-Image-2512-GGUF), \[Qwen-Image-Edit-2511\](https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF), and \[Qwen-Image-Layered\](https://huggingface.co/unsloth/Qwen-Image-Layered-GGUF) {% hint style="warning" %} The format of the vae and diffusion model might be different than the diffusers checkpoints if using checkpoints other than the ones above. Only use checkpoints that are compatible with ComfyUI. {% endhint %} These files must be in the correct folders for ComfyUI to see them. In addition the vision tower stored in the mmproj file must use the same prefix as the text encoder. Download reference images to be used later as well: \`\`\`bash curl -L -C - -o ../input/sloth1.jpg \\ "https://unsloth.ai/cgi/image/\_1d5a5685-2d88-44ca-b50f-ba432cd646ef\_9CGCY8lvw4D9JkOdueqsk.jpeg?width=1920&quality=80&format=jpeg" curl -L -C - -o ../input/sloth2.jpg \\ "https://unsloth.ai/cgi/image/UnSloth\_GPU\_Front\_-\_Confetti\_ArcSk-MR4MMN215UutOFZ.png?width=1920&quality=80&format=jpeg" \`\`\` ### #3. Workflow and Hyperparameters For more info you can also view our detailed \[Run GGUFs in ComfyUI\](/docs/blog/comfyui.md#workflow-and-hyperparameters-1) Guide. Navigate to the main ComfyUI directory and run: \`\`\`bash python main.py \`\`\` {% hint style="info" %} \`python main.py --cpu\` to run with CPU, but will be slow. {% endhint %} This will launch a web server that allows you to access \`https://127.0.0.1:8188\` . If you are running this on the cloud, you'll need to make sure port forwarding is setup to access on your local machine. Workflows are saved as JSON files embedded in output images (PNG metadata) or as separate \`.json\` files. You can: \* Drag & drop an image into ComfyUI to load its workflow \* Export/import workflows via the menu \* Share workflows as JSON files Below are two examples of Qwen-Image-2512 and Qwen-Image-Edit-2511 json files which you can download and use: {% file src="/files/GEtfTCPAWM4rFYgNvYfd" %} For our workflow, we default to \*\*1024×1024\*\* as a practical middle ground. While the model supports native resolution (1328×1328), generating at native typically increases runtime by \*\*\\~50%\*\*. Since GGUF adds overhead and 40 steps is a relatively long run, 1024×1024 keeps generation time reasonable. If needed, you can increase resolution to 1328. {% hint style="warning" %} For more realistic results, skip keywords like “photorealistic” or “digital rendering” or “3d render” and use terms like “photograph” instead. {% endhint %} {% hint style="info" %} For negative prompts, it’s best to use an NLP-style approach: describe in \*\*natural language\*\* what you \*don’t\* want in the image. Packing in too many keywords can hurt results instead of making it more specific. {% endhint %} {% file src="/files/Z0JySsOLKNNUZb6tzGDk" %} {% columns %} {% column %} Instead of setting up the workflow from scratch you can download the workflow here. Load it into the browser page by clicking the Comfy Logo -> File -> Open -> Then choose the \`unsloth\_qwen\_image\_2512.json\` file you just downloaded. It should look like the below: {% endcolumn %} {% column %} ![](https://unsloth.ai/files/H9Gjy6PMp1YOarv773sG) {% endcolumn %} {% endcolumns %} ![](https://unsloth.ai/files/QOis9p890bpMO2l6rFCc) This workflow is based on the official ComfyUI published workflow except it uses the GGUF loader extension, and is simplified to illustrate text to image functionality. ### #4. Inference ComfyUI is highly customizable. You can mix models and create extremely complex pipelines. For a basic text to image setup we need to load the model, specify prompt and image details, and decide on a sampling strategy. #### \*\*Upload Models + Set Prompt\*\* We already downloaded the models, so we just need to pick the correct ones. For Unet Loader pick \`qwen-image-2512-Q4\_K\_M.gguf\`, for CLIPLoader pick \`Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf\`, and for Load VAE pick \`qwen\_image\_vae.safetensors\`. {% hint style="info" %} For more realistic results, skip keywords like “photorealistic” or “digital rendering” or “3d render” and use terms like “photograph” instead. {% endhint %} You can set any prompt you'd like, and also specify a negative prompt. The negative prompt helps by telling the model where to steer away from. {% hint style="info" %} For negative prompts, it’s best to use an NLP-style approach: describe in \*\*natural language\*\* what you \*don’t\* want in the image. Packing in too many keywords can hurt results instead of making it more specific. {% endhint %} #### \*\*Image Size + Sampler Parameters\*\* The Qwen Image model series supports different image sizes. You can make rectangular shapes by setting the values of width and height. For sampler parameters, you can experiment with different samplers other than euler, and more or less sampling steps. The workflow has steps set to 40, but for quick tests 20 might be good enough. Change the \`control after generate\` setting from randomize to fixed if you want to see how different settings change outputs. #### \*\*Run\*\* Click Run and an image will be generated in about 1 minute (30 seconds for 20 steps). That output image can be saved. The interesting part is that the metadata for the entire comfy workflow is saved in the image. You can share and anyone can see how it was created by loading it in the UI. ![](https://unsloth.ai/files/qDZCkONxtdjvFhSaYy8Z) {% hint style="info" %} If you're encountering blurry/bad images, raise shift to 12-13! solves most issues with bad outputs. {% endhint %} #### \*\*Multi Reference Generation\*\* A key feature of Qwen-Image-Edit-2511 is multi reference generation where you can supply multiple images to use to help control generation. This time load the \`unsloth\_qwen\_image\_edit\_2511.json\`. We will use most of the same models but switching \`qwen-image-2512-Q4\_K\_M.gguf\` to \`qwen-image-edit-2511-Q4\_K\_M.gguf\` for the unet. The other difference this time are extra nodes to select images to reference, which we've downloaded earlier. You'll notice the prompt refers to both \`image 1\` and \`image 2\` which are prompt anchors for the images. Once loaded click Run, and you'll see an output that creates our two unique sloth characters together while preserving their likeness. ![](https://unsloth.ai/files/tlhXPSWN4S4l8XGFybab) Final result made from images to the right: ![](https://unsloth.ai/files/dinNIgwE83v9z27KTbcm) ![](https://unsloth.ai/files/wWLCjKRl0jjQSDM5KEPU) \## 🤗 D\*\*iffusers Tutorial\*\* We have also uploaded a \[Dynamic 4-bit BitsandBytes\](https://huggingface.co/unsloth/Qwen-Image-2512-unsloth-bnb-4bit) quantized version which can be run in Hugging Face's \`diffusers\` library. Once again, it uses Unsloth Dynamic where important layers are upcasted to higher precision. Run \`Qwen-Image-2512-unsloth-bnb-4bit\` with the code below: \`\`\`python from diffusers import DiffusionPipeline import torch pipe = DiffusionPipeline.from\_pretrained( "unsloth/Qwen-Image-2512-unsloth-bnb-4bit", torch\_dtype=torch.bfloat16, ).to('cuda') # uncomment if you run out of memory # pipe.enable\_model\_cpu\_offload() output = pipe( prompt="a kawaii sloth playing the drums", negative\_prompt="blurry, unfocused", num\_inference\_steps=20, true\_cfg\_scale=4.0, ) # Save output image = output.images\[0\] image.save('sample.png') \`\`\` ## 🎨 \*\*stable-diffusion.cpp Tutorial\*\* If you want to run the model in stable-diffusion.cpp, you can follow our \[step-by-step guide here\](/docs/models/tutorials/qwen-image-2512/stable-diffusion.cpp.md). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/qwen-image-2512.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/minimax-m27.md). # MiniMax-M2.7 - How to Run Locally MiniMax-M2.7 is a new open model for agentic coding and chat use-cases. The model achieves SOTA performance in SWE-Pro (56.22%) and Terminal Bench 2 (57.0%). The \*\*230B parameters\*\* (10B active) model is the successor to \[MiniMax-M25\](/docs/models/tutorials/minimax-m25.md) and has a \*\*200K context\*\* window. The unquantized bf16 requires \*\*457GB\*\*. Unsloth Dynamic \*\*4-bit\*\* GGUF reduces the size to \*\*108GB\*\* \*\*(-60%)\*\* so it can run on a \*\*128GB RAM\*\* device\*\*:\*\* \[\*\*MiniMax-M2.7 GGUF\*\*\](https://huggingface.co/unsloth/MiniMax-M2.7-GGUF) All uploads use Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) for SOTA quantization performance - so important layers are upcasted to higher bits (e.g. 8 or 16-bit). Thank you MiniMax for day zero access. {% hint style="success" %} NEW MiniMax-M2.7 GGUF Benchmarks available! \[See here\](#gguf-benchmarks) {% endhint %} ### :gear: Usage Guide The 4-bit dynamic quant \`UD-IQ4\_XS\` uses \*\*108GB\*\* of disk space - this fits nicely on a \*\*128GB unified memory Mac\*\* for \\~15+ tokens/s, and also works faster with a \*\*1x16GB GPU and 96GB of RAM\*\* for 25+ tokens/s. \*\*2-bit\*\* quants or the biggest 2-bit will fit on a 96GB device. For near \*\*full precision\*\*, use \`Q8\_0\` (8-bit) which utilizes 243GB and will fit on a 256GB RAM device / Mac for 15+ tokens/s. {% hint style="success" %} For best performance, make sure your total available memory (VRAM + system RAM) exceeds the size of the quantized model file you’re downloading. If it doesn’t, llama.cpp can still run via SSD/HDD offloading, but inference will be slower. {% endhint %} ### Recommended Settings MiniMax recommends using the following parameters for best performance: \`temperature=1.0\`, \`top\_p = 0.95\`, \`top\_k = 40\`. {% columns %} {% column %} | Default Settings (Most Tasks) | | ----------------------------- | | \`temperature = 1.0\` | | \`top\_p = 0.95\` | | \`top\_k = 40\` | | {% endcolumn %} | {% column %} \* \*\*Maximum context window:\*\* \`196,608\` \* Default system prompt: {% code overflow="wrap" %} \`\`\` You are a helpful assistant. Your name is MiniMax-M2.7 and is built by MiniMax. \`\`\` {% endcode %} {% endcolumn %} {% endcolumns %} ## Run MiniMax-M2.7 Tutorials: To make MiniMax-M2.7 work on a 128GB RAM device, we will be utilizing the 4-bit \[\`UD-IQ4\_XS\` quant\](https://huggingface.co/unsloth/MiniMax-M2.7-GGUF?show\_file\_info=UD-IQ4\_XS%2FMiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf). You can now run MiniMax-M2.7 in \[llama.cpp\](#run-in-llama.cpp) and \[Unsloth Studio\](#run-in-unsloth-studio). {% hint style="warning" %} Do NOT use CUDA 13.2 to run any model as it may cause gibberish or poor outputs. NVIDIA is working on a fix. {% endhint %} ### 🦥 Run in Unsloth Studio MiniMax-M2.7 can now run in \[Unsloth Studio\](/docs/new/studio.md), our new open-source web UI for local AI. Unsloth Studio lets you run models locally on \*\*MacOS, Windows\*\*, Linux and: {% columns %} {% column %} \* Search, download, \[run GGUFs\](/docs/new/studio.md#run-models-locally) and safetensor models \* \[\*\*Self-healing\*\* tool calling\](/docs/new/studio.md#execute-code--heal-tool-calling) + \*\*web search\*\* \* \[\*\*Code execution\*\*\](/docs/new/studio.md#run-models-locally) (Python, Bash) \* \[Automatic inference\](/docs/new/studio.md#model-arena) parameter tuning (temp, top-p, etc.) \* Uses llama.cpp for Fast CPU + GPU inference and CPU offloading {% endcolumn %} {% column %} ![](https://unsloth.ai/files/3rrq58PcvZFnywcYbnPb) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Install Unsloth Run in your terminal: \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% endstep %} {% step %} #### Launch Unsloth \*\*MacOS, Linux, WSL and Windows:\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Then open \`http://localhost:8888\` in your browser.\*\* {% endstep %} {% step %} #### Search and download MiniMax-M2.7 On first launch you will need to create a password to secure your account and sign in again later. You’ll then see a brief onboarding wizard to choose a model, dataset, and basic settings. You can skip it at any time. You can choose \`UD-IQ4\_XS\` (dynamic 4bit quant) or other quantized versions like \`UD-Q4\_K\_XL\` . If downloads get stuck, see \[Hugging Face Hub, XET debugging\](/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) Then go to the \[Unsloth Chat\](/docs/new/studio/chat.md) tab and search for MiniMax-M2.7 in the search bar and download your desired model and quant. It will take some time to download due to the size so please wait. To ensure fast inference, ensure you have \[enough RAM/VRAM\](#usage-guide), otherwise inference will still work, but Unsloth will offload to your CPU. ![](https://unsloth.ai/files/2vifoqs4j1dOnyCPyWhj) {% endstep %} {% step %} #### Run MiniMax-M2.7 Inference parameters should be auto-set when using Unsloth Studio, however you can still change it manually. You can also edit the context length, chat template and other settings. For more information, you can view our \[Unsloth Studio inference guide\](/docs/new/studio/chat.md). {% endstep %} {% endstepper %} ### ✨ Run in llama.cpp {% hint style="warning" %} Do NOT use CUDA 13.2 to run any model as it may cause gibberish or poor outputs. NVIDIA is working on a fix. {% endhint %} {% stepper %} {% step %} Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} If you want to use \`llama.cpp\` directly to load models, you can do the below: (:IQ4\\\_XS) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. Remember the model has only a maximum of 200K context length. Follow this for \*\*most default\*\* use-cases: \`\`\`bash export LLAMA\_CACHE="unsloth/MiniMax-M2.7-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/MiniMax-M2.7-GGUF:UD-IQ4\_XS \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 40 \`\`\` {% endstep %} {% step %} Download the model (after installing \`pip install huggingface\_hub hf\_transfer\`). You can choose UD-IQ4\\\_XS (dynamic 4-bit quant) or other quantized versions like \`UD-Q6\_K\_XL\` . We recommend using our 4bit dynamic quant UD-IQ4\\\_XS to balance size and accuracy. If downloads get stuck, see \[Hugging Face Hub, XET debugging\](/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) \`\`\`bash hf download unsloth/MiniMax-M2.7-GGUF \\ --local-dir unsloth/MiniMax-M2.7-GGUF \\ --include "\*UD-IQ4\_XS\*" # Use "\*Q8\_0\*" for 8-bit \`\`\` {% endstep %} {% step %} You can edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length, \`--n-gpu-layers 2\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/MiniMax-M2.7-GGUF/UD-IQ4\_XS/MiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 40 \`\`\` {% endcode %} {% endstep %} {% endstepper %} #### 🦙 Llama-server & OpenAI's completion library To deploy MiniMax-M2.7 for production, we use \`llama-server\` or OpenAI API. In a new terminal say via tmux, deploy the model via: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\ --model unsloth/MiniMax-M2.7-GGUF/UD-IQ4\_XS/MiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf \\ --alias "unsloth/MiniMax-M2.7" \\ --prio 3 \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.01 \\ --top-k 40 \\ --port 8001 \`\`\` {% endcode %} Then in a new terminal, after doing \`pip install openai\`, do: {% code overflow="wrap" %} \`\`\`python from openai import OpenAI import json openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/MiniMax-M2.7", messages = \[{"role": "user", "content": "Create a Snake game."},\], ) print(completion.choices\[0\].message.content) \`\`\` {% endcode %} ## 📊 Benchmarks ### GGUF Benchmarks Below are KLD 99% benchmarks for MiniMax-M2.7. Lower left is better: ![](https://unsloth.ai/files/rIP8x2x6o3j2Zot3xiKC) Because MiniMax-M2.7 utilizes the same architecture as MiniMax-M2.5, GGUF quantization benchmarks for M2.7 should be very similar to M2.5. So, we'll also refer to previous quant benchmark conducted for M2.5: ![](https://unsloth.ai/files/hfUzL4ykvVI3HWR95ySj) \[Benjamin Marie (third-party) benchmarked\](https://x.com/bnjmn\_marie/status/2027043753484021810/photo/1) \*\*MiniMax-M2.5\*\* using \*\*Unsloth GGUF quantizations\*\* on a \*\*750-prompt mixed suite\*\* (LiveCodeBench v6, MMLU Pro, GPQA, Math500), reporting both \*\*overall accuracy\*\* and \*\*relative error increase\*\* (how much more often the quantized model makes mistakes vs. the original). Unsloth quants, no matter their precision perform much better than their non-Unsloth counterparts for both accuracy and relative error (despite being 8GB smaller). \*\*Key results:\*\* \* \*\*Best quality/size tradeoff here: \`unsloth UD-Q4\_K\_XL\`.\*\*\\ It’s the closest to Original: only \*\*6.0 points\*\* down, and “only” \*\*+22.8%\*\* more errors than baseline. \* \*\*Other Unsloth Q4 quants perform closely together (\\~64.5–64.9 accuracy).\*\*\\ \`IQ4\_NL\`, \`MXFP4\_MOE\`, and \`UD-IQ2\_XXS\` are all basically the same quality on this benchmark, with \*\*\\~33–35%\*\* more errors than Original. \* Unsloth GGUFs perform much better than other non-Unsloth GGUFs, e.g. see \`lmstudio-community - Q4\_K\_M\` (despite being 8GB smaller) and \`AesSedai - IQ3\_S\`. ### Official Benchmarks ![](https://unsloth.ai/files/0QvHO0YnxQ7W9b08wujn) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/minimax-m27.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/ministral-3.md). # Ministral 3 - How to Run Guide Mistral releases Ministral 3, their new multimodal models in Base, Instruct, and Reasoning variants, available in \*\*3B\*\*, \*\*8B\*\*, and \*\*14B\*\* sizes. They offer best-in-class performance for their size, and are fine-tuned for instruction and chat use cases. The multimodal models support \*\*256K context\*\* windows, multiple languages, native function calling, and JSON output. The full unquantized 14B Ministral-3-Instruct-2512 model fits in \*\*24GB RAM\*\*/VRAM. You can now run, fine-tune and RL on all Ministral 3 models with Unsloth: [Run Ministral 3 Tutorials](https://unsloth.ai/docs/models/tutorials/ministral-3.md#run-ministral-3-tutorials) [Fine-tuning Ministral 3](https://unsloth.ai/pages/zRbjQXuLmfdZD90U410W#fine-tuning) We've also uploaded Mistral Large 3 \[GGUFs here\](https://huggingface.co/unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF). For all Ministral 3 uploads (BnB, FP8), \[see here\](https://huggingface.co/collections/unsloth/ministral-3). | Ministral-3-Instruct GGUFs: | Ministral-3-Reasoning GGUFs: | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | \[3B\](https://huggingface.co/unsloth/Ministral-3-3B-Instruct-2512-GGUF) • \[8B\](https://huggingface.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF) • \[14B\](https://huggingface.co/unsloth/Ministral-3-14B-Instruct-2512-GGUF) | \[3B\](https://huggingface.co/unsloth/Ministral-3-3B-Reasoning-2512-GGUF) • \[8B\](https://huggingface.co/unsloth/Ministral-3-8B-Reasoning-2512-GGUF) • \[14B\](https://huggingface.co/unsloth/Ministral-3-14B-Reasoning-2512-GGUF) | ### ⚙️ Usage Guide To achieve optimal performance for \*\*Instruct\*\*, Mistral recommends using lower temperatures such as \`temperature = 0.15\` or \`0.1\` For \*\*Reasoning\*\*, Mistral recommends \`temperature = 0.7\` and \`top\_p = 0.95\`. | Instruct: | Reasoning: | | ----------------------------- | ------------------- | | \`Temperature = 0.15\` or \`0.1\` | \`Temperature = 0.7\` | | \`Top\_P = default\` | \`Top\_P = 0.95\` | \*\*Adequate Output Length\*\*: Use an output length of \`32,768\` tokens for most queries for the reasoning variant, and \`16,384\` for the instruct variant. You can increase the max output size for the reasoning model if necessary. The maximum context length Ministral 3 can reach is \`262,144\` The chat template format is found when we use the below: {% code overflow="wrap" %} \`\`\`python tokenizer.apply\_chat\_template(\[ {"role" : "user", "content" : "What is 1+1?"}, {"role" : "assistant", "content" : "2"}, {"role" : "user", "content" : "What is 2+2?"} \], add\_generation\_prompt = True ) \`\`\` {% endcode %} #### Ministral \*Reasoning\* chat template: {% code overflow="wrap" lineNumbers="true" %} \`\`\` ~\[SYSTEM\_PROMPT\]# HOW YOU SHOULD THINK AND ANSWER First draft your thinking process (inner monologue) until you arrive at a response. Format your response using Markdown, and use LaTeX for any mathematical equations. Write both your thoughts and the response in the same language as the input. Your thinking process must follow the template below:\[THINK\]Your thoughts or/and draft, like working through an exercise on scratch paper. Be as casual and as long as you want until you are confident to generate the response to the user.\[/THINK\]Here, provide a self-contained response.\[/SYSTEM\_PROMPT\]\[INST\]What is 1+1?\[/INST\]2~\[INST\]What is 2+2?\[/INST\] \`\`\` {% endcode %} #### Ministral \*Instruct\* chat template: {% code overflow="wrap" lineNumbers="true" expandable="true" %} \`\`\` ~\[SYSTEM\_PROMPT\]You are Ministral-3-3B-Instruct-2512, a Large Language Model (LLM) created by Mistral AI, a French startup headquartered in Paris. You power an AI assistant called Le Chat. Your knowledge base was last updated on 2023-10-01. The current date is {today}. When you're not sure about some information or when the user's request requires up-to-date or specific data, you must use the available tools to fetch the information. Do not hesitate to use tools whenever they can provide a more accurate or complete response. If no relevant tools are available, then clearly state that you don't have the information and avoid making up anything. If the user's question is not clear, ambiguous, or does not provide enough context for you to accurately answer the question, you do not try to answer it right away and you rather ask the user to clarify their request (e.g. "What are some good restaurants around me?" => "Where are you?" or "When is the next flight to Tokyo" => "Where do you travel from?"). You are always very attentive to dates, in particular you try to resolve dates (e.g. "yesterday" is {yesterday}) and when asked about information at specific dates, you discard information that is at another date. You follow these instructions in all languages, and always respond to the user in the language they use or request. Next sections describe the capabilities that you have. # WEB BROWSING INSTRUCTIONS You cannot perform any web search or access internet to open URLs, links etc. If it seems like the user is expecting you to do so, you clarify the situation and ask the user to copy paste the text directly in the chat. # MULTI-MODAL INSTRUCTIONS You have the ability to read images, but you cannot generate images. You also cannot transcribe audio files or videos. You cannot read nor transcribe audio files or videos. # TOOL CALLING INSTRUCTIONS You may have access to tools that you can use to fetch information or perform actions. You must use these tools in the following situations: 1. When the request requires up-to-date information. 2. When the request requires specific data that you do not have in your knowledge base. 3. When the request involves actions that you cannot perform without tools. Always prioritize using tools to provide the most accurate and helpful response. If tools are not available, inform the user that you cannot perform the requested action at the moment.\[/SYSTEM\_PROMPT\]\[INST\]What is 1+1?\[/INST\]2~\[INST\]What is 2+2?\[/INST\] \`\`\` {% endcode %} ## 📖 Run Ministral 3 Tutorials Below are guides for the \[Reasoning\](#reasoning-ministral-3-reasoning-2512) and \[Instruct\](#instruct-ministral-3-instruct-2512) variants of the model. ### Instruct: Ministral-3-Instruct-2512 To achieve optimal performance for \*\*Instruct\*\*, Mistral recommends using lower temperatures such as \`temperature = 0.15\` or \`0.1\` #### :sparkles: Llama.cpp: Run Ministral-3-14B-Instruct Tutorial {% stepper %} {% step %} Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} You can directly pull from Hugging Face via: \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Ministral-3-14B-Instruct-2512-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32784 \\ --temp 0.15 \`\`\` {% endstep %} {% step %} Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD\_Q4\_K\_XL\` or other quantized versions. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Ministral-3-14B-Instruct-2512-GGUF", local\_dir = "Ministral-3-14B-Instruct-2512-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` {% endstep %} {% endstepper %} ### Reasoning: Ministral-3-Reasoning-2512 To achieve optimal performance for \*\*Reasoning\*\*, Mistral recommends using \`temperature = 0.7\` and \`top\_p = 0.95\`. #### :sparkles: Llama.cpp: Run Ministral-3-14B-Reasoning Tutorial {% stepper %} {% step %} Obtain the latest \`llama.cpp\` on \[GitHub\](https://github.com/ggml-org/llama.cpp). You can also use the build instructions below. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} You can directly pull from Hugging Face via: \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Ministral-3-14B-Reasoning-2512-GGUF:Q4\_K\_XL \\ --jinja -ngl 99 --ctx-size 32784 \\ --temp 0.6 --top-p 0.95 \`\`\` {% endstep %} {% step %} Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD\_Q4\_K\_XL\` or other quantized versions. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Ministral-3-14B-Reasoning-2512-GGUF", local\_dir = "Ministral-3-14B-Reasoning-2512-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` {% endstep %} {% endstepper %} ## 🛠️ Fine-tuning Ministral 3 [](https://unsloth.ai/docs/models/tutorials/ministral-3.md#fine-tuning) Unsloth now supports fine-tuning of all Ministral 3 models, including vision support. To train, you must use the latest 🤗Hugging Face \`transformers\` v5 and \`unsloth\` which includes our our recent \[ultra long context\](/docs/blog/500k-context-length-fine-tuning.md) support. The large 14B Ministral 3 model should fit on a free Colab GPU. We made free Unsloth notebooks to fine-tune Ministral 3. Change the name to use the desired model. \* Ministral-3B-Instruct \[Vision notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_VL\_\\(3B\\)\_Vision.ipynb) (vision) \* Ministral-3B-Instruct \[GRPO notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) {% columns %} {% column %} Ministral Vision finetuning notebook {% embed url="" %} {% endcolumn %} {% column %} Ministral Sudoku GRPO RL notebook {% embed url="" %} {% endcolumn %} {% endcolumns %} ### :sparkles:Reinforcement Learning (GRPO) Unsloth now supports RL and GRPO for the Mistral models as well. As usual, they benefit from all of Unsloth's enhancements and tomorrow, we are going to release a notebook soon specifically for autonomously solving the sudoku puzzle. \* Ministral-3B-Instruct \[GRPO notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Ministral\_3\_\\(3B\\)\_Reinforcement\_Learning\_Sudoku\_Game.ipynb) \*\*To use the latest version of Unsloth and transformers v5, update via:\*\* {% code overflow="wrap" %} \`\`\` pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth\_zoo \`\`\` {% endcode %} The goal is to auto generate strategies to complete Sudoku! {% columns %} {% column %} ![](https://unsloth.ai/files/dCO8Qa6Wvp7EhmdwmG8P) {% endcolumn %} {% column %} ![](https://unsloth.ai/files/tLr39wHu1ZcHOFkAkqFN) {% endcolumn %} {% endcolumns %} For the reward plots for Ministral, we get the below. We see it works well! {% columns %} {% column %} !\[\](/files/maqkwoZcOrDGxK277M6c) !\[\](/files/AJGnL7GAIZBWZHLhn7KH) {% endcolumn %} {% column %} !\[\](/files/bAFJ2KDXHxaJSRksoemE) !\[\](/files/yXtqDtsugRGvdDGOLBAm) {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/ministral-3.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/devstral-how-to-run-and-fine-tune.md). # Devstral : comment l'exécuter et le fine-tuner \*\*Devstral-Small-2507\*\* (Devstral 1.1) est le nouveau LLM agentique de Mistral pour l’ingénierie logicielle. Il excelle dans l’appel d’outils, l’exploration de bases de code et l’alimentation d’agents de codage. Mistral AI a publié la version originale 2505 en mai 2025. Affiné à partir de \[\*\*Mistral-Small-3.1\*\*\](https://huggingface.co/unsloth/Mistral-Small-3.1-24B-Instruct-2503-GGUF), Devstral prend en charge une fenêtre de contexte de 128k. Devstral Small 1.1 a de meilleures performances, atteignant un score de 53,6 % sur \[SWE-bench vérifié\](https://openai.com/index/introducing-swe-bench-verified/), ce qui en fait (10 juillet 2025) le modèle ouvert n°1 sur le benchmark. Les GGUF Unsloth Devstral 1.1 contiennent des \*\*prise en charge de l’appel d’outils\*\* et \*\*corrections du modèle de chat\*\*. Devstral 1.1 fonctionne toujours bien avec OpenHands, mais généralise désormais mieux à d’autres invites et environnements de codage. En mode texte uniquement, l’encodeur de vision de Devstral a été retiré avant l’affinage. Nous avons ajouté \[\*\*\*une prise en charge Vision optionnelle\*\*\*\](#possible-vision-support) pour le modèle. {% hint style="success" %} Nous avons également travaillé en coulisses avec Mistral pour aider à déboguer, tester et corriger d’éventuels bugs et problèmes ! Assurez-vous de \*\*télécharger les téléchargements officiels de Mistral ou les GGUF d’Unsloth\*\* / les quants dynamiques pour obtenir \*\*l’implémentation correcte\*\* (c.-à-d. le bon prompt système, le bon modèle de chat, etc.) Veuillez utiliser \`--jinja\` dans llama.cpp pour activer le prompt système ! {% endhint %} Tous les envois Devstral utilisent notre \[Dynamic 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) méthodologie Unsloth, offrant les meilleures performances sur les benchmarks 5-shot MMLU et KL Divergence. Cela signifie que vous pouvez exécuter et affiner des LLM Mistral quantifiés avec une perte d’exactitude minimale ! #### \*\*Devstral - Unsloth Dynamic\*\* quants : | Devstral 2507 (nouveau) | Devstral 2505 | | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | GGUF : \[Devstral-Small-2507-GGUF\](https://huggingface.co/unsloth/Devstral-Small-2507-GGUF) | \[Devstral-Small-2505-GGUF\](https://huggingface.co/unsloth/Devstral-Small-2505-GGUF) | | 4-bit BnB : \[Devstral-Small-2507-unsloth-bnb-4bit\](https://huggingface.co/unsloth/Devstral-Small-2507-unsloth-bnb-4bit) | \[Devstral-Small-2505-unsloth-bnb-4bit\](https://huggingface.co/unsloth/Devstral-Small-2505-unsloth-bnb-4bit) | ## 🖥️ \*\*Exécution de Devstral\*\* ### :gear: Paramètres recommandés officiels Selon Mistral AI, voici les paramètres recommandés pour l’inférence : \* \*\*Température de 0,0 à 0,15\*\* \* Min\\\_P de 0,01 (optionnel, mais 0,01 fonctionne bien, la valeur par défaut de llama.cpp est 0,1) \* \*\*Utilisez\*\*\*\* \*\*\*\*\`--jinja\`\*\*\*\* \*\*\*\*pour activer le prompt système.\*\* \*\*Un prompt système est recommandé\*\*, et dérive du prompt système d’OpenHands. Le prompt système complet est fourni \[ici\](https://huggingface.co/unsloth/Devstral-Small-2505/blob/main/SYSTEM\_PROMPT.txt). \`\`\` Vous êtes Devstral, un modèle agentique utile entraîné par Mistral AI et utilisant l’ossature OpenHands. Vous pouvez interagir avec un ordinateur pour résoudre des tâches. Votre rôle principal est d’assister les utilisateurs en exécutant des commandes, en modifiant du code et en résolvant efficacement des problèmes techniques. Vous devez être rigoureux, méthodique et privilégier la qualité plutôt que la vitesse. \* Si l’utilisateur pose une question, comme « pourquoi X se produit-il », n’essayez pas de corriger le problème. Répondez simplement à la question. .... LE PROMPT SYSTÈME CONTINUE .... \`\`\` {% hint style="success" %} Nos envois dynamiques ont le préfixe '\`UD\`'. Ceux qui n’en ont pas ne sont pas dynamiques, mais utilisent tout de même notre jeu de données de calibration. {% endhint %} ## :llama: Tutoriel : comment exécuter Devstral dans Ollama 1. Installer \`ollama\` si vous ne l’avez pas déjà fait ! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Exécutez le modèle avec notre quantification dynamique. Notez que vous pouvez appeler \`ollama serve &\`dans un autre terminal si cela échoue ! Nous incluons tous les paramètres suggérés (température, etc.) dans \`params\` dans notre envoi Hugging Face ! 3. De plus, Devstral prend en charge des longueurs de contexte de 128K, il est donc préférable d’activer \[\*\*la quantification du cache KV\*\*\](https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-set-the-quantization-type-for-the-kv-cache). Nous utilisons une quantification 8 bits qui économise 50 % d’utilisation mémoire. Vous pouvez aussi essayer \`"q4\_0"\` \`\`\`bash export OLLAMA\_KV\_CACHE\_TYPE="q8\_0" ollama run hf.co/unsloth/Devstral-Small-2507-GGUF:UD-Q4\_K\_XL \`\`\` ## 📖 Tutoriel : comment exécuter Devstral dans llama.cpp 1. Obtenez le dernier \`llama.cpp\` par défaut. Seule votre machine peut atteindre le serveur. \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` à \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d’habitude - la prise en charge Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\\\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Si vous voulez utiliser \`llama.cpp\` directement pour charger des modèles, vous pouvez faire ce qui suit : (:Q4\\\_K\\\_XL) est le type de quantification. Vous pouvez également télécharger via Hugging Face (point 3). C’est similaire à \`ollama run\` \`\`\`bash ./llama.cpp/llama-cli -hf unsloth/Devstral-Small-2507-GGUF:UD-Q4\_K\_XL --jinja \`\`\` 3. \*\*OU\*\* télécharger le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir Q4\\\_K\\\_M, ou d’autres versions quantifiées (comme BF16 en précision complète). \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Devstral-Small-2507-GGUF", local\_dir = "unsloth/Devstral-Small-2507-GGUF", allow\_patterns = \["\*Q4\_K\_XL\*", "\*mmproj-F16\*"\], # Pour Q4\_K\_XL ) \`\`\` 4. Exécutez le modèle. 5. Modifiez \`--threads -1\` pour le nombre maximal de threads CPU, \`--ctx-size 131072\` pour la longueur du contexte (Devstral prend en charge une longueur de contexte de 128K !), \`--n-gpu-layers 99\` pour le déchargement GPU sur le nombre de couches. Essayez de l’ajuster si votre GPU manque de mémoire. Supprimez-le également si vous n’avez qu’une inférence CPU. Nous utilisons aussi une quantification 8 bits pour le cache K afin de réduire l’utilisation mémoire. 6. En mode conversation : ./llama.cpp/llama-cli \\ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4_K_XL.gguf \\ --threads -1 \\ --ctx-size 131072 \\ --cache-type-k q8_0 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 0.15 \\ --repeat-penalty 1.0 \\ --min-p 0.01 \\ --top-k 64 \\ --top-p 0.95 \\ --jinja 7\. Pour le mode non conversationnel afin de tester notre invite Flappy Bird : ./llama.cpp/llama-cli \\ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4_K_XL.gguf \\ --threads -1 \\ --ctx-size 131072 \\ --cache-type-k q8_0 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\ --temp 0.15 \\ --repeat-penalty 1.0 \\ --min-p 0.01 \\ --top-k 64 \\ --top-p 0.95 \\ -no-cnv \\ --prompt "[SYSTEM_PROMPT]Vous êtes Devstral, un modèle agentique utile entraîné par Mistral AI et utilisant l’ossature OpenHands. Vous pouvez interagir avec un ordinateur pour résoudre des tâches.\n\n\nVotre rôle principal est d’assister les utilisateurs en exécutant des commandes, en modifiant du code et en résolvant efficacement des problèmes techniques. Vous devez être rigoureux, méthodique et privilégier la qualité plutôt que la vitesse.\n* Si l’utilisateur pose une question, comme \"pourquoi X se produit-il\", n’essayez pas de corriger le problème. Répondez simplement à la question.\n\n\n\n* Chaque action que vous entreprenez a un certain coût. Dans la mesure du possible, combinez plusieurs actions en une seule, par exemple en regroupant plusieurs commandes bash en une seule, en utilisant sed et grep pour modifier/afficher plusieurs fichiers à la fois.\n* Lors de l’exploration de la base de code, utilisez des outils efficaces comme find, grep et les commandes git avec les filtres appropriés afin de minimiser les opérations inutiles.\n\n\n\n* Lorsqu’un utilisateur fournit un chemin de fichier, ne supposez PAS qu’il est relatif au répertoire de travail actuel. Explorez d’abord le système de fichiers pour localiser le fichier avant d’y travailler.\n* Si l’on vous demande de modifier un fichier, modifiez le fichier directement plutôt que d’en créer un nouveau avec un nom différent.\n* Pour les opérations globales de recherche-remplacement, envisagez d’utiliser `sed` plutôt que d’ouvrir plusieurs fois des éditeurs de fichiers.\n\n\n\n* Écrivez un code propre et efficace avec un minimum de commentaires. Évitez la redondance dans les commentaires : ne répétez pas des informations qui peuvent être facilement déduites du code lui-même.\n* Lorsque vous implémentez des solutions, concentrez-vous sur les modifications minimales nécessaires pour résoudre le problème.\n* Avant d’implémenter des changements, comprenez d’abord en profondeur la base de code par l’exploration.\n* Si vous ajoutez beaucoup de code à une fonction ou à un fichier, envisagez de diviser la fonction ou le fichier en parties plus petites lorsque cela est approprié.\n\n\n\n* Lors de la configuration des identifiants git, utilisez par défaut \"openhands\" comme user.name et \"openhands@all-hands.dev\" comme user.email, sauf instruction contraire explicite.\n* Faites preuve de prudence avec les opérations git. Ne faites PAS de modifications potentiellement dangereuses (par exemple pousser vers main, supprimer des dépôts) sauf si cela est explicitement demandé.\n* Lors de la validation des changements, utilisez `git status` pour voir tous les fichiers modifiés et préparez tous les fichiers nécessaires au commit. Utilisez `git commit -a` chaque fois que possible.\n* Ne validez PAS de fichiers qui ne devraient généralement pas aller dans le contrôle de version (par exemple node_modules/, fichiers .env, répertoires de build, fichiers cache, gros binaires) sauf instruction explicite de l’utilisateur.\n* En cas de doute sur certains fichiers à valider, vérifiez la présence de fichiers .gitignore ou demandez des précisions à l’utilisateur.\n\n\n\n* Lors de la création de demandes d’extraction, créez-en une seule par session/problème, sauf instruction explicite contraire.\n* Lorsque vous travaillez avec une PR existante, mettez-la à jour avec de nouveaux commits plutôt que de créer des PR supplémentaires pour le même problème.\n* Lors de la mise à jour d’une PR, conservez le titre et l’objectif d’origine, en mettant à jour la description uniquement si nécessaire.\n\n\n\n1. EXPLORATION : explorez minutieusement les fichiers pertinents et comprenez le contexte avant de proposer des solutions\n2. ANALYSIS : considérez plusieurs approches et sélectionnez la plus prometteuse\n3. TESTING :\n * Pour les corrections de bugs : créez des tests pour vérifier les problèmes avant d’implémenter les corrections\n * Pour les nouvelles fonctionnalités : envisagez le développement piloté par les tests lorsque cela est approprié\n * Si le dépôt ne dispose pas d’une infrastructure de tests et que la mise en place des tests nécessiterait une configuration importante, consultez l’utilisateur avant d’investir du temps dans la construction de cette infrastructure\n * Si l’environnement n’est pas configuré pour exécuter des tests, consultez d’abord l’utilisateur avant d’investir du temps pour installer toutes les dépendances\n4. IMPLEMENTATION : apportez des modifications ciblées et minimales pour résoudre le problème\n5. VERIFICATION : si l’environnement est configuré pour exécuter des tests, testez votre implémentation de manière approfondie, y compris les cas limites. Si l’environnement n’est pas configuré pour exécuter des tests, consultez d’abord l’utilisateur avant d’investir du temps dans l’exécution des tests.\n\n\n\n* N’utilisez les identifiants GITHUB_TOKEN et autres identifiants que de la manière explicitement demandée et attendue par l’utilisateur.\n* Utilisez des API pour travailler avec GitHub ou d’autres plateformes, sauf si l’utilisateur demande autrement ou si votre tâche nécessite la navigation.\n\n\n\n* Lorsque l’utilisateur vous demande d’exécuter une application, ne vous arrêtez pas si l’application n’est pas installée. Installez plutôt l’application et relancez la commande.\n* Si vous rencontrez des dépendances manquantes :\n 1. Cherchez d’abord dans le dépôt des fichiers de dépendances existants (requirements.txt, pyproject.toml, package.json, Gemfile, etc.)\n 2. Si des fichiers de dépendances existent, utilisez-les pour installer toutes les dépendances en une fois (par exemple, `pip install -r requirements.txt`, `npm install`, etc.)\n 3. N’installez directement des paquets individuels que si aucun fichier de dépendances n’est trouvé ou si seuls des paquets spécifiques sont nécessaires\n* De même, si vous rencontrez des dépendances manquantes pour des outils essentiels demandés par l’utilisateur, installez-les lorsque c’est possible.\n\n\n\n* Si vous avez fait plusieurs tentatives pour résoudre un problème mais que les tests échouent toujours ou que l’utilisateur indique que ce n’est toujours pas corrigé :\n 1. Reculer et réfléchir à 5 à 7 sources possibles différentes du problème\n 2. Évaluer la probabilité de chacune de ces causes possibles\n 3. Traiter méthodiquement les causes les plus probables, en commençant par la probabilité la plus élevée\n 4. Documenter votre raisonnement\n* Lorsque vous rencontrez un problème majeur en exécutant un plan de l’utilisateur, n’essayez pas de le contourner directement. Proposez plutôt un nouveau plan et obtenez la confirmation de l’utilisateur avant de poursuivre.\n[/SYSTEM_PROMPT][INST]Créez un jeu Flappy Bird en Python. Vous devez inclure ces éléments :\n1. Vous devez utiliser pygame.\n2. La couleur de fond doit être choisie aléatoirement et être une teinte claire. Commencez avec une couleur bleu clair.\n3. Appuyer plusieurs fois sur SPACE accélérera l’oiseau.\n4. La forme de l’oiseau doit être choisie aléatoirement entre un carré, un cercle ou un triangle. La couleur doit être choisie aléatoirement et être sombre.\n5. Placez en bas un sol de couleur brun foncé ou jaune, choisi aléatoirement.\n6. Affichez un score en haut à droite. Incrémentez-le si vous passez les tuyaux sans les heurter.\n7. Faites des tuyaux espacés aléatoirement avec suffisamment d’espace. La couleur doit être choisie aléatoirement parmi un vert foncé, un marron clair ou une nuance de gris foncé.\n8. Quand vous perdez, affichez le meilleur score. Le texte doit être à l’intérieur de l’écran. Appuyer sur q ou Esc quittera le jeu. Redémarrer se fait en appuyant à nouveau sur SPACE.\nLe jeu final doit être inclus dans une section markdown en Python. Vérifiez votre code pour détecter les erreurs[/INST]" {% hint style="danger" %} N’oubliez pas de supprimer \\ puisque Devstral ajoute automatiquement un \\ ! Utilisez aussi \`--jinja\` pour activer le prompt système ! {% endhint %} ## :eyes:Prise en charge expérimentale de la vision \[Xuan-Son\](https://x.com/ngxson) de Hugging Face a montré dans leur \[dépôt GGUF\](https://huggingface.co/ngxson/Devstral-Small-Vision-2505-GGUF) qu’il est en réalité possible de « greffer » l’encodeur de vision de Mistral 3.1 Instruct sur Devstral 2507. Nous avons également téléversé nos fichiers mmproj, ce qui vous permet d’utiliser ce qui suit : \`\`\`bash ./llama.cpp/llama-mtmd-cli \\\\ --model unsloth/Devstral-Small-2507-GGUF/Devstral-Small-2507-UD-Q4\_K\_XL.gguf \\\\ --mmproj unsloth/Devstral-Small-2507-GGUF/mmproj-F16.gguf \\\\ --threads -1 \\\\ --ctx-size 131072 \\\\ --cache-type-k q8\_0 \\\\ --n-gpu-layers 99 \\\\ --seed 3407 \\\\ --prio 2 \\\\ --temp 0.15 \`\`\` Par exemple : | Code d’instruction et de sortie | Code rendu | | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | !\[\](https://cdn-uploads.huggingface.co/production/uploads/63ca214abedad7e2bf1d1517/HDic53ANsCoJbiWu2eE6K.png) | !\[\](https://cdn-uploads.huggingface.co/production/uploads/63ca214abedad7e2bf1d1517/onV1xfJIT8gzh81RkLn8J.png) | ## 🦥 Affinage de Devstral avec Unsloth Tout comme les modèles Mistral standard, y compris Mistral Small 3.1, Unsloth prend en charge l’affinage de Devstral. L’entraînement est 2 fois plus rapide, utilise 70 % de VRAM en moins et prend en charge des longueurs de contexte 8 fois plus longues. Devstral tient confortablement dans un GPU L4 de 24 Go de VRAM. Malheureusement, Devstral dépasse légèrement les limites de mémoire d’une VRAM de 16 Go, donc l’affiner gratuitement sur Google Colab n’est pas possible pour le moment. Cependant, vous \*pouvez\* affiner le modèle gratuitement en utilisant notre \[notebook Kaggle\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Magistral\_\\(24B\\)-Reasoning-Conversational.ipynb\\&accelerator=nvidiaTeslaT4), qui offre l’accès à deux GPU. Il suffit de modifier le nom du modèle Magistral du notebook pour le modèle Devstral. Si vous avez une ancienne version d’Unsloth et/ou si vous affineez localement, installez la dernière version d’Unsloth : \`\`\`bash pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` \[^1\]: Quantification K pour réduire l’utilisation mémoire. Peut être f16, q8\\\_0, q4\\\_0 \[^2\]: Doit utiliser --jinja pour activer le prompt système --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/devstral-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/how-to-run-llms-with-docker.md). # Comment exécuter des LLM locaux avec Docker : guide étape par étape Vous pouvez désormais exécuter n'importe quel modèle, y compris Unsloth \[GGUF dynamiques\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md), sur Mac, Windows ou Linux avec une seule ligne de code ou \*\*aucun code\*\* du tout. Nous avons collaboré avec Docker pour simplifier le déploiement des modèles, et Unsloth alimente désormais la plupart des modèles GGUF sur Docker. Avant de commencer, assurez-vous de consulter \[exigences matérielles\](#hardware-info--performance) et \[nos conseils\](#hardware-info--performance) pour optimiser les performances lors de l'exécution de LLM sur votre appareil. [Tutoriel Docker Terminal](https://unsloth.ai/pages/18ff82c8ec22aac46f659c98c536562dad45be0b#method-1-docker-terminal) [Tutoriel Docker sans code](https://unsloth.ai/docs/fr/modeles/tutorials/how-to-run-llms-with-docker.md#method-2-docker-desktop-no-code) Pour commencer, exécutez OpenAI \[gpt-oss\](/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune.md) avec une seule commande : \`\`\`bash docker model run ai/gpt-oss:20B \`\`\` Ou pour exécuter un \[modèle Unsloth\](/docs/fr/commencer/unsloth-model-catalog.md) / quant depuis Hugging Face : \`\`\`bash docker model run hf.co/unsloth/gpt-oss-20b-GGUF:F16 \`\`\` {% hint style="success" %} Vous n'avez pas besoin de Docker Desktop, Docker CE suffit pour exécuter les modèles. {% endhint %} #### \*\*Pourquoi Unsloth + Docker ?\*\* Nous collaborons avec des labs de modèles comme Google Gemma pour corriger les bugs des modèles et améliorer la précision. Nos GGUF dynamiques surpassent systématiquement les autres méthodes de quantification, vous offrant une inférence précise et efficace. Si vous utilisez Docker, vous pouvez exécuter des modèles instantanément sans configuration. Docker utilise \[Docker Model Runner\](https://github.com/docker/model-runner) (DMR), qui vous permet d'exécuter des LLM aussi facilement que des conteneurs sans problèmes de dépendances. DMR utilise les modèles Unsloth et \`llama.cpp\` sous le capot pour une inférence rapide, efficace et à jour. ## :gear: Infos Matériel + Performance Pour de meilleures performances, visez à ce que votre VRAM + RAM combinées soient au moins égales à la taille du modèle quantifié que vous téléchargez. Si vous en avez moins, le modèle fonctionnera toujours, mais beaucoup plus lentement. Assurez-vous également que votre appareil dispose de suffisamment d'espace disque pour stocker le modèle. Si votre modèle tient à peine en mémoire, vous pouvez vous attendre à environ \\~5 tokens/s, selon la taille du modèle. Disposer de RAM/VRAM supplémentaire améliorera la vitesse d'inférence, et une VRAM additionnelle permettra le plus grand gain de performances (à condition que l'ensemble du modèle tienne) {% hint style="info" %} \*\*Exemple :\*\* Si vous téléchargez gpt-oss-20b (F16) et que le modèle fait 13,8 Go, assurez-vous que votre espace disque et votre RAM + VRAM > 13,8 Go. {% endhint %} \*\*Recommandations de quantification :\*\* \* Pour les modèles de moins de 30 milliards de paramètres, utilisez au moins 4 bits (Q4). \* Pour les modèles de 70 milliards de paramètres ou plus, utilisez un minimum de quantification 2 bits (par ex., UD\\\_Q2\\\_K\\\_XL). ## ⚡ Tutoriels pas à pas Ci-dessous se trouvent \*\*deux façons\*\* d'exécuter des modèles avec Docker : l'une en utilisant le \[terminal\](#method-1-docker-terminal), et l'autre en utilisant \[Docker Desktop\](#method-2-docker-desktop-no-code) sans code : ### Méthode n°1 : Docker Terminal {% stepper %} {% step %} #### Installer Docker Docker Model Runner est déjà disponible dans \*\*les deux\*\* \[Docker Desktop\](https://docs.docker.com/ai/model-runner/get-started/#docker-desktop) et \[\*\*Docker CE\*\*\](https://docs.docker.com/ai/model-runner/get-started/#docker-engine)\*\*.\*\* {% endstep %} {% step %} #### Exécuter le modèle Choisissez un modèle à exécuter, puis lancez la commande via le terminal. \* Parcourez le catalogue vérifié des modèles de confiance disponibles sur \[Docker Hub\](https://hub.docker.com/r/ai) ou \[La page Hugging Face d'Unsloth\](https://huggingface.co/unsloth) . \* Allez dans le Terminal pour exécuter les commandes. Pour vérifier si vous avez \`docker\` installé, vous pouvez taper 'docker' et appuyer sur Entrée. \* Docker Hub lance par défaut Unsloth Dynamic 4-bit, cependant vous pouvez choisir votre propre niveau de quantification (voir l'étape n°3). Par exemple, pour exécuter OpenAI \`gpt-oss-20b\` en une seule commande : \`\`\`bash docker model run ai/gpt-oss:20B \`\`\` Ou pour exécuter un \[Unsloth\](/docs/fr/commencer/unsloth-model-catalog.md) gpt-oss quant depuis Hugging Face : \`\`\`bash docker model run hf.co/unsloth/gpt-oss-20b-GGUF:UD-Q8\_K\_XL \`\`\` \*\*Voici à quoi devrait ressembler l'exécution de gpt-oss-20b via CLI :\*\* ![](https://unsloth.ai/files/2166e664b4eae5fa45eedfbea1c9ad066a3ff27a) gpt-oss-20b depuis Docker Hub ![](https://unsloth.ai/files/b499bde85c313bfb9170d79954f1728bbe303ee4) gpt-oss-20b avec la quantification UD-Q8\_K\_XL d'Unsloth {% endstep %} {% step %} #### Pour exécuter un niveau de quantification spécifique : Si vous souhaitez exécuter une quantification spécifique d'un modèle, ajoutez \`:\` et le nom de la quantification au modèle (par ex., \`Q4\` pour Docker ou \`UD-Q4\_K\_XL\`). Vous pouvez voir toutes les quantifications disponibles sur la page Docker Hub de chaque modèle. par ex. voir les quantifications listées pour gpt-oss \[ici\](https://hub.docker.com/r/ai/gpt-oss#gptoss). La même chose s'applique aux quants Unsloth sur Hugging Face : visitez la \[page HF du modèle\](https://huggingface.co/unsloth/gpt-oss-20b-GGUF?show\_file\_info=gpt-oss-20b-Q2\_K\_L.gguf), choisissez une quantification, puis exécutez quelque chose comme : \`docker model run hf.co/unsloth/gpt-oss-20b-GGUF:Q2\_K\_L\` ![](https://unsloth.ai/files/0fec9ef1cf522bab4867efd64a38b5dfe51cabe9) Niveaux de quantification gpt-oss sur [Docker Hub](https://hub.docker.com/r/ai/gpt-oss#gptoss) ![](https://unsloth.ai/files/f4365df8b1572fb4eb536b1e295f6bf4084c55a8) Niveaux de quantification Unsloth gpt-oss sur [Hugging Face](https://huggingface.co/unsloth/gpt-oss-20b-GGUF) {% endstep %} {% endstepper %} ### Méthode n°2 : Docker Desktop (sans code) {% stepper %} {% step %} #### Installer Docker Desktop Docker Model Runner est déjà disponible dans \[Docker Desktop\](https://docs.docker.com/ai/model-runner/get-started/#docker-desktop). 1. Choisissez un modèle à exécuter, ouvrez Docker Desktop, puis cliquez sur l'onglet modèles. 2. Cliquez sur 'Add models +' ou Docker Hub. Recherchez le modèle. Parcourez le catalogue de modèles vérifiés disponible sur \[Docker Hub\](https://hub.docker.com/r/ai). ![](https://unsloth.ai/files/a78b03e8f5909b031bf2896425d874e739590e60) #1. Cliquez sur l'onglet 'Models' puis sur 'Add models +' ![](https://unsloth.ai/files/ae2ceaebd8014f256c88a8995f1aa76f27242d80) #2. Recherchez le modèle souhaité. {% endstep %} {% step %} #### Télécharger le modèle Cliquez sur le modèle que vous souhaitez exécuter pour voir les quantifications disponibles. \* Les quantifications vont de 1 à 16 bits. Pour les modèles de moins de 30 milliards de paramètres, utilisez au moins 4 bits (\`Q4\`). \* Choisissez une taille qui correspond à votre matériel : idéalement, votre mémoire unifiée combinée, RAM ou VRAM devrait être égale ou supérieure à la taille du modèle. Par exemple, un modèle de 11 Go fonctionne bien sur 12 Go de mémoire unifiée. ![](https://unsloth.ai/files/c669fee1d59b4aa574374c71597a8a8e1a251e90) #3. Sélectionnez la quantification que vous souhaitez télécharger. ![](https://unsloth.ai/files/8d4a1c0d667403592f66c1c7f3ba4037c9bafa23) #4. Attendez que le modèle ait fini de se télécharger, puis exécutez-le. {% endstep %} {% step %} #### Exécuter le modèle Tapez n'importe quelle invite dans la case 'Ask a question' et utilisez le LLM comme vous utiliseriez ChatGPT. ![](https://unsloth.ai/files/f98c4b745f2a78cac22247e51b58aa72a7e89e31) Un exemple d'exécution de Qwen3-4B `UD-Q8_K_XL` {% endstep %} {% endstepper %} #### \*\*Pour exécuter les modèles les plus récents :\*\* Vous pouvez exécuter n'importe quel nouveau modèle sur Docker tant qu'il est pris en charge par \`llama.cpp\` ou \`vllm\` et disponible sur Docker Hub. ### Qu'est-ce que Docker Model Runner ? Le Docker Model Runner (DMR) est un outil open source qui vous permet de télécharger et d'exécuter des modèles d'IA aussi facilement que vous exécutez des conteneurs. GitHub : Il fournit un runtime cohérent pour les modèles, similaire à la façon dont Docker a standardisé le déploiement d'applications. Sous le capot, il utilise des backends optimisés (comme \`llama.cpp\`) pour une inférence fluide et efficace en ressources sur votre machine. Que vous soyez chercheur, développeur ou amateur, vous pouvez désormais : \* Exécuter des modèles ouverts localement en quelques secondes. \* Éviter l'enfer des dépendances, tout est géré dans Docker. \* Partager et reproduire des configurations de modèles sans effort. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/how-to-run-llms-with-docker.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/qwen-image-2512.md). # Comment exécuter Qwen-Image-2512 localement dans ComfyUI \*\*Qwen-Image-2512\*\* est la mise à jour de décembre des modèles fondamentaux de Qwen pour la génération d'images à partir de texte. Le modèle est le modèle de diffusion open-source le plus performant et ce guide vous apprendra à l'exécuter localement via \[Unsloth\](https://github.com/unslothai/unsloth) GGUF et ComfyUI. Qwen-Image-2512 fonctionnalités : des personnes au rendu plus réaliste ; des détails plus riches dans les paysages/textures ; et un rendu du texte plus précis. \*\*Téléversements :\*\* \[GGUF\](https://huggingface.co/unsloth/Qwen-Image-2512-GGUF) • \[FP8\](https://huggingface.co/unsloth/Qwen-Image-2512-FP8) • \[BnB 4 bits\](https://huggingface.co/unsloth/Qwen-Image-2512-unsloth-bnb-4bit) Les quantifications utilisent \[Unsloth Dynamic\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) méthodologie qui relève certaines couches importantes à une précision supérieure pour récupérer plus de précision. Merci à Qwen d'avoir permis le support Unsloth dès le jour 0. ## 📖 Tutoriel ComfyUI Pour l'exécuter, vous n'avez pas besoin d'un GPU, un CPU avec de la RAM suffit. Pour de meilleurs résultats, assurez-vous que votre mémoire totale utilisable (RAM + VRAM / unifiée) est supérieure à la taille GGUF ; par exemple 4 bits (Q4\\\_K\\\_M) \`unsloth/Qwen-Image-Edit-2512-GGUF\` fait 13,1 Go, donc vous devriez avoir 13,2+ Go de mémoire combinée. \[ComfyUI\](https://github.com/Comfy-Org/ComfyUI) est une interface graphique open-source, une API et un back-end pour modèles de diffusion qui utilise une interface basée sur des nœuds (graphe/organigramme). Ce guide se concentrera sur les machines avec CUDA, mais les instructions pour construire sur Apple ou CPU sont similaires. ### #1. Installation & Configuration Pour installer ComfyUI, vous pouvez télécharger l'application de bureau sur les appareils Windows ou Mac \[ici\](https://www.comfy.org/download). Sinon, pour configurer ComfyUI afin d'exécuter des modèles GGUF, exécutez ce qui suit : \`\`\`bash mkdir comfy\_ggufs cd comfy\_ggufs python -m venv .venv source .venv/bin/activate git clone https://github.com/Comfy-Org/ComfyUI.git cd ComfyUI pip install -r requirements.txt cd custom\_nodes git clone https://github.com/city96/ComfyUI-GGUF cd ComfyUI-GGUF pip install -r requirements.txt cd ../.. \`\`\` ### #2. Télécharger les modèles Les modèles de diffusion nécessitent généralement 3 modèles. Un Variational AutoEncoder (VAE) qui encode l'espace pixel de l'image en espace latent, un encodeur de texte pour traduire le texte en embeddings d'entrée, et le transformeur de diffusion proprement dit. Vous pouvez trouver tous les GGUF de diffusion Unsloth dans notre \[Collection ici\](https://huggingface.co/collections/unsloth/unsloth-diffusion-ggufs). Le modèle de diffusion et l'encodeur de texte peuvent être au format GGUF tandis que nous utilisons généralement safetensors pour le VAE. Selon \[le dépôt de Qwen\](https://huggingface.co/Qwen/Qwen-Image-2512/blob/main/text\_encoder/config.json), nous utiliserons Qwen2.5-VL et non \[Qwen3-VL\](/docs/fr/modeles/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md). Téléchargeons les modèles que nous utiliserons (Remarque : vous pouvez aussi utiliser notre \[téléversement FP8\](https://huggingface.co/unsloth/Qwen-Image-2512-FP8) dans ComfyUI) : \`\`\`bash cd models ## Modèles de diffusion curl -L -C - -o unet/qwen-image-2512-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-2512-GGUF/resolve/main/qwen-image-2512-Q4\_K\_M.gguf curl -L -C - -o unet/qwen-image-edit-2511-Q4\_K\_M.gguf \\ https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF/resolve/main/qwen-image-edit-2511-Q4\_K\_M.gguf ## Encodeur de texte + Vision Tower + VAE curl -L -C - -o text\_encoders/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf \\ https://huggingface.co/unsloth/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/Qwen2.5-VL-7B-Instruct-UD-Q4\_K\_XL.gguf curl -L -C - -o text\_encoders/Qwen2.5-VL-7B-Instruct-mmproj-BF16.gguf \\ https://huggingface.co/unsloth/Qwen2.5-VL-7B-Instruct-GGUF/resolve/main/mmproj-BF16.gguf curl -L -C - -o vae/qwen\_image\_vae.safetensors \\ https://huggingface.co/Comfy-Org/Qwen-Image\_ComfyUI/resolve/main/split\_files/vae/qwen\_image\_vae.safetensors \`\`\` Voir les téléversements GGUF pour : \[Qwen-Image-2512\](https://huggingface.co/unsloth/Qwen-Image-2512-GGUF), \[Qwen-Image-Edit-2511\](https://huggingface.co/unsloth/Qwen-Image-Edit-2511-GGUF), et \[Qwen-Image-Layered\](https://huggingface.co/unsloth/Qwen-Image-Layered-GGUF) {% hint style="warning" %} Le format du VAE et du modèle de diffusion peut être différent des checkpoints diffusers si vous utilisez des checkpoints autres que ceux ci-dessus. N'utilisez que des checkpoints compatibles avec ComfyUI. {% endhint %} Ces fichiers doivent être dans les dossiers corrects pour que ComfyUI puisse les voir. De plus, la vision tower stockée dans le fichier mmproj doit utiliser le même préfixe que l'encodeur de texte. Téléchargez également des images de référence qui seront utilisées plus tard : \`\`\`bash curl -L -C - -o ../input/sloth1.jpg \\ "https://unsloth.ai/cgi/image/\_1d5a5685-2d88-44ca-b50f-ba432cd646ef\_9CGCY8lvw4D9JkOdueqsk.jpeg?width=1920&quality=80&format=jpeg" curl -L -C - -o ../input/sloth2.jpg \\ "https://unsloth.ai/cgi/image/UnSloth\_GPU\_Front\_-\_Confetti\_ArcSk-MR4MMN215UutOFZ.png?width=1920&quality=80&format=jpeg" \`\`\` ### #3. Workflow et hyperparamètres Pour plus d'infos vous pouvez également consulter notre \[Run GGUFs in ComfyUI\](/docs/fr/blog/comfyui.md#workflow-and-hyperparameters-1) Guide. Allez dans le répertoire principal de ComfyUI et exécutez : \`\`\`bash python main.py \`\`\` {% hint style="info" %} \`python main.py --cpu\` pour exécuter avec le CPU, mais ce sera lent. {% endhint %} Cela lancera un serveur web qui vous permettra d'accéder à \`https://127.0.0.1:8188\` . Si vous exécutez cela sur le cloud, vous devrez vous assurer que le transfert de port est configuré pour y accéder depuis votre machine locale. Les workflows sont sauvegardés en tant que fichiers JSON intégrés dans les images de sortie (métadonnées PNG) ou en tant que \`.json\` fichiers. Vous pouvez : \* Glisser-déposer une image dans ComfyUI pour charger son workflow \* Exporter/importer des workflows via le menu \* Partager des workflows sous forme de fichiers JSON Ci-dessous se trouvent deux exemples de fichiers json pour Qwen-Image-2512 et Qwen-Image-Edit-2511 que vous pouvez télécharger et utiliser : {% file src="/files/eac5939aba3b1e281d77ce90ed5b4825ef10408e" %} Pour notre workflow, nous utilisons par défaut \*\*1024×1024\*\* comme compromis pratique. Bien que le modèle prenne en charge la résolution native (1328×1328), générer en natif augmente généralement le temps d'exécution de \*\*\\~50%\*\*. Puisque GGUF ajoute une surcharge et 40 étapes est une exécution relativement longue, 1024×1024 maintient un temps de génération raisonnable. Si nécessaire, vous pouvez augmenter la résolution à 1328. {% hint style="warning" %} Pour des résultats plus réalistes, évitez des mots-clés comme « photoréaliste » ou « rendu numérique » ou « rendu 3D » et utilisez plutôt des termes comme « photographie ». {% endhint %} {% hint style="info" %} Pour les prompts négatifs, il est préférable d'utiliser une approche de type PNL : décrivez en \*\*langage naturel\*\* ce que \*vous ne\* souhaitez pas {% endhint %} {% file src="/files/27049e4a7739518e73b7e484d31da05c2968e8b8" %} {% columns %} {% column %} dans l'image. Empiler trop de mots-clés peut nuire aux résultats au lieu de les rendre plus spécifiques. Au lieu de configurer le workflow depuis zéro, vous pouvez télécharger le workflow ici. \`Chargez-le dans la page du navigateur en cliquant sur le logo Comfy -> Fichier -> Ouvrir -> Puis choisissez le\` unsloth\\\_qwen\\\_image\\\_2512.json {% endcolumn %} {% column %} ![](https://unsloth.ai/files/811d5fa42b28949cd23a0f9d3986d099db85a7e5) {% endcolumn %} {% endcolumns %} ![](https://unsloth.ai/files/3486aaa9f47ebd2fb78699ea2cb41af4819d1f16) fichier que vous venez de télécharger. Il devrait ressembler à ce qui suit : ### Ce workflow est basé sur le workflow officiel publié par ComfyUI sauf qu'il utilise l'extension de chargement GGUF, et il est simplifié pour illustrer la fonctionnalité de texte vers image. \\#4. Inférence #### \*\*ComfyUI est hautement personnalisable. Vous pouvez mélanger des modèles et créer des pipelines extrêmement complexes. Pour une configuration basique de texte vers image, nous devons charger le modèle, spécifier le prompt et les détails d'image, et décider d'une stratégie d'échantillonnage.\*\* Téléverser les modèles + Définir le prompt \`Nous avons déjà téléchargé les modèles, donc nous devons juste choisir les bons. Pour Unet Loader choisissez\`qwen-image-2512-Q4\\\_K\\\_M.gguf \`, pour CLIPLoader choisissez\`Qwen2.5-VL-7B-Instruct-UD-Q4\\\_K\\\_XL.gguf \`, et pour Load VAE choisissez\`. {% hint style="info" %} Pour des résultats plus réalistes, évitez des mots-clés comme « photoréaliste » ou « rendu numérique » ou « rendu 3D » et utilisez plutôt des termes comme « photographie ». {% endhint %} qwen\\\_image\\\_vae.safetensors {% hint style="info" %} Pour les prompts négatifs, il est préférable d'utiliser une approche de type PNL : décrivez en \*\*langage naturel\*\* ce que \*vous ne\* souhaitez pas {% endhint %} #### \*\*Vous pouvez définir n'importe quel prompt que vous souhaitez, et aussi spécifier un prompt négatif. Le prompt négatif aide en indiquant au modèle ce dont il doit s'éloigner.\*\* Taille d'image + paramètres du sampler \`La série Qwen Image prend en charge différentes tailles d'image. Vous pouvez créer des formes rectangulaires en réglant les valeurs de largeur et hauteur. Pour les paramètres du sampler, vous pouvez expérimenter avec des samplers différents de euler, et plus ou moins d'étapes d'échantillonnage. Le workflow a les étapes réglées à 40, mais pour des tests rapides 20 peut suffire. Changez le\` contrôle après génération #### \*\*paramètre de randomize à fixed si vous voulez voir comment différents réglages modifient les sorties.\*\* Exécuter ![](https://unsloth.ai/files/41f164e1f03d3995dccb8b13eb7ca1a897df5c91) {% hint style="info" %} Cliquez sur Exécuter et une image sera générée en environ 1 minute (30 secondes pour 20 étapes). Cette image de sortie peut être sauvegardée. La partie intéressante est que les métadonnées de l'ensemble du workflow comfy sont sauvegardées dans l'image. Vous pouvez partager et n'importe qui peut voir comment elle a été créée en la chargeant dans l'UI. {% endhint %} #### \*\*Si vous rencontrez des images floues/mauvaises, augmentez shift à 12-13 ! résout la plupart des problèmes d'images de mauvaise qualité.\*\* Génération multi-références \`Une fonctionnalité clé de Qwen-Image-Edit-2511 est la génération multi-références où vous pouvez fournir plusieurs images pour aider à contrôler la génération. Cette fois chargez le\`unsloth\\\_qwen\\\_image\\\_edit\\\_2511.json \`Nous avons déjà téléchargé les modèles, donc nous devons juste choisir les bons. Pour Unet Loader choisissez\` . Nous utiliserons la plupart des mêmes modèles mais en changeant \`pour\` qwen-image-edit-2511-Q4\\\_K\\\_M.gguf \`pour l'unet. L'autre différence cette fois est l'ajout de nœuds supplémentaires pour sélectionner les images de référence, que nous avons téléchargées plus tôt. Vous remarquerez que le prompt fait référence à la fois à\` image 1 \`et\` image 2 ![](https://unsloth.ai/files/67a7315ca4ee177c70a55b9a514b2ef2b1721d2e) qui sont des ancres de prompt pour les images. Une fois chargé, cliquez sur Exécuter, et vous verrez une sortie qui crée nos deux personnages paresseux uniques ensemble tout en préservant leur ressemblance. ![](https://unsloth.ai/files/c63bb48a0f61f744394218357c69f97385962183) ![](https://unsloth.ai/files/ef05a4dc1332418a77d4245cc9eebe8f2cc1a1d3) \## Résultat final réalisé à partir des images à droite :\*\*🤗 D\*\* iffusers Tutoriel \[Nous avons également téléversé une\](https://huggingface.co/unsloth/Qwen-Image-2512-unsloth-bnb-4bit) version quantifiée Dynamic 4-bit BitsandBytes \`qui peut être exécutée avec la bibliothèque\` diffusers paramètre de randomize à fixed si vous voulez voir comment différents réglages modifient les sorties. \`de Hugging Face. Encore une fois, elle utilise Unsloth Dynamic où les couches importantes sont relevées à une précision supérieure.\` Qwen-Image-2512-unsloth-bnb-4bit \`\`\`python avec le code ci-dessous : from diffusers import DiffusionPipeline import torch pipe = DiffusionPipeline.from\_pretrained( "unsloth/Qwen-Image-2512-unsloth-bnb-4bit", torch\_dtype=torch.bfloat16, ).to('cuda') # décommentez si vous manquez de mémoire # pipe.enable\_model\_cpu\_offload() output = pipe( prompt="un paresseux kawaii jouant de la batterie", negative\_prompt="flou, hors de focus", num\_inference\_steps=20, ) true\_cfg\_scale=4.0, # Sauvegarder la sortie image = output.images\[0\] \`\`\` ## image.save('sample.png') \*\*🎨\*\* Tutoriel stable-diffusion.cpp \[Si vous souhaitez exécuter le modèle dans stable-diffusion.cpp, vous pouvez suivre notre guide pas à pas ici\](/docs/fr/modeles/tutorials/qwen-image-2512/stable-diffusion.cpp.md). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/qwen-image-2512.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/phi-4-reasoning-how-to-run-and-fine-tune.md). # Phi-4 Reasoning: How to Run & Fine-tune Microsoft's new Phi-4 reasoning models are now supported in Unsloth. The 'plus' variant performs on par with OpenAI's o1-mini, o3-mini and Sonnet 3.7. The 'plus' and standard reasoning models are 14B parameters while the 'mini' has 4B parameters.\\ \\ All Phi-4 reasoning uploads use our \[Unsloth Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) methodology. #### \*\*Phi-4 reasoning - Unsloth Dynamic 2.0 uploads:\*\* | Dynamic 2.0 GGUF (to run) | Dynamic 4-bit Safetensor (to finetune/deploy) | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | * [Reasoning-plus](https://huggingface.co/unsloth/Phi-4-reasoning-plus-GGUF/) (14B) * [Reasoning](https://huggingface.co/unsloth/Phi-4-reasoning-GGUF) (14B) * [Mini-reasoning](https://huggingface.co/unsloth/Phi-4-mini-reasoning-GGUF/) (4B) | * [Reasoning-plus](https://huggingface.co/unsloth/Phi-4-reasoning-plus-unsloth-bnb-4bit) * [Reasoning](https://huggingface.co/unsloth/phi-4-reasoning-unsloth-bnb-4bit) * [Mini-reasoning](https://huggingface.co/unsloth/Phi-4-mini-reasoning-unsloth-bnb-4bit) | ## 🖥️ \*\*Running Phi-4 reasoning\*\* ### :gear: Official Recommended Settings According to Microsoft, these are the recommended settings for inference: \* \*\*Temperature = 0.8\*\* \* Top\\\_P = 0.95 ### \*\*Phi-4 reasoning Chat templates\*\* Please ensure you use the correct chat template as the 'mini' variant has a different one. #### \*\*Phi-4-mini:\*\* {% code overflow="wrap" %} \`\`\` <|system|>Your name is Phi, an AI math expert developed by Microsoft.<|end|><|user|>How to solve 3\*x^2+4\*x+5=1?<|end|><|assistant|> \`\`\` {% endcode %} #### \*\*Phi-4-reasoning and Phi-4-reasoning-plus:\*\* This format is used for general conversation and instructions: {% code overflow="wrap" %} \`\`\` <|im\_start|>system<|im\_sep|>You are Phi, a language model trained by Microsoft to help users. Your role as an assistant involves thoroughly exploring questions through a systematic thinking process before providing the final precise and accurate solutions. This requires engaging in a comprehensive cycle of analysis, summarizing, exploration, reassessment, reflection, backtracing, and iteration to develop well-considered thinking process. Please structure your response into two main sections: Thought and Solution using the specified format: {Thought section} {Solution section}. In the Thought section, detail your reasoning process in steps. Each step should include detailed considerations such as analysing questions, summarizing relevant findings, brainstorming new ideas, verifying the accuracy of the current steps, refining any errors, and revisiting previous steps. In the Solution section, based on various attempts, explorations, and reflections from the Thought section, systematically present the final solution that you deem correct. The Solution section should be logical, accurate, and concise and detail necessary steps needed to reach the conclusion. Now, try to solve the following question through the above guidelines:<|im\_end|><|im\_start|>user<|im\_sep|>What is 1+1?<|im\_end|><|im\_start|>assistant<|im\_sep|> \`\`\` {% endcode %} {% hint style="info" %} Yes, the chat template/prompt format is this long! {% endhint %} ### 🦙 Ollama: Run Phi-4 reasoning Tutorial 1. Install \`ollama\` if you haven't already! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Run the model! Note you can call \`ollama serve\`in another terminal if it fails. We include all our fixes and suggested parameters (temperature etc) in \`params\` in our Hugging Face upload. \`\`\`bash ollama run hf.co/unsloth/Phi-4-mini-reasoning-GGUF:Q4\_K\_XL \`\`\` ### 📖 Llama.cpp: Run Phi-4 reasoning Tutorial {% hint style="warning" %} You must use \`--jinja\` in llama.cpp to enable reasoning for the models, expect for the 'mini' variant. Otherwise no token will be provided. {% endhint %} 1. Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose Q4\\\_K\\\_M, or other quantized versions. \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Phi-4-mini-reasoning-GGUF", local\_dir = "unsloth/Phi-4-mini-reasoning-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], ) \`\`\` 3. Run the model in conversational mode in llama.cpp. You must use \`--jinja\` in llama.cpp to enable reasoning for the models. This is however not needed if you're using the 'mini' variant. \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Phi-4-mini-reasoning-GGUF/Phi-4-mini-reasoning-UD-Q4\_K\_XL.gguf \\ --threads -1 \\ --n-gpu-layers 99 \\ --prio 3 \\ --temp 0.8 \\ --top-p 0.95 \\ --jinja \\ --min-p 0.00 \\ --ctx-size 32768 \\ --seed 3407 \`\`\` ## 🦥 Fine-tuning Phi-4 with Unsloth \[Phi-4 fine-tuning\](https://unsloth.ai/blog/phi4) for the models are also now supported in Unsloth. To fine-tune for free on Google Colab, just change the \`model\_name\` of 'unsloth/Phi-4' to 'unsloth/Phi-4-mini-reasoning' etc. \* \[Phi-4 (14B) fine-tuning notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/phi-4-reasoning-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/minimax-m27.md). # MiniMax-M2.7 - Comment l'exécuter en local MiniMax-M2.7 est un nouveau modèle ouvert pour les cas d’usage de codage agentique et de chat. Le modèle atteint des performances SOTA dans SWE-Pro (56,22 %) et Terminal Bench 2 (57,0 %). Le \*\*230B paramètres\*\* (10B actifs) le modèle est le successeur de \[MiniMax-M25\](/docs/fr/modeles/tutorials/minimax-m25.md) et a une \*\*fenêtre de contexte de 200K\*\* fenêtre. Le bf16 non quantifié nécessite \*\*457 Go\*\*. Unsloth Dynamic \*\*4 bits\*\* GGUF réduit la taille à \*\*108 Go\*\* \*\*(-60%)\*\* afin qu’il puisse fonctionner sur un \*\*128 Go de RAM\*\* appareil\*\*:\*\* \[\*\*GGUF MiniMax-M2.7\*\*\](https://huggingface.co/unsloth/MiniMax-M2.7-GGUF) Tous les téléversements utilisent Unsloth \[Dynamique 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) pour des performances de quantification SOTA — ainsi les couches importantes sont remontées vers des bits plus élevés (par ex. 8 ou 16 bits). Merci à MiniMax pour l’accès dès le jour zéro. {% hint style="success" %} Nouveaux benchmarks GGUF de MiniMax-M2.7 disponibles ! \[Voir ici\](#gguf-benchmarks) {% endhint %} ### :gear: Guide d'utilisation La quantification dynamique 4 bits \`UD-IQ4\_XS\` utilise \*\*108 Go\*\* d’espace disque — cela tient parfaitement sur un \*\*Mac avec 128 Go de mémoire unifiée\*\* pour \\~15+ jetons/s, et fonctionne aussi plus vite avec un \*\*GPU 1x16 Go et 96 Go de RAM\*\* pour 25+ jetons/s. \*\*2 bits\*\* les quants ou le plus gros modèle 2 bits tiendront sur un appareil de 96 Go. Pour une performance proche de \*\*précision complète\*\*, utilisez \`Q8\_0\` (8 bits) qui utilise 243 Go et tiendra sur un appareil / Mac avec 256 Go de RAM pour 15+ jetons/s. {% hint style="success" %} Pour de meilleures performances, assurez-vous que votre mémoire totale disponible (VRAM + RAM système) dépasse la taille du fichier de modèle quantifié que vous téléchargez. Si ce n'est pas le cas, llama.cpp peut quand même fonctionner via un déchargement SSD/HDD, mais l'inférence sera plus lente. {% endhint %} ### Paramètres recommandés MiniMax recommande d’utiliser les paramètres suivants pour de meilleures performances : \`temperature=1.0\`, \`top\_p = 0,95\`, \`top\_k = 40\`. {% columns %} {% column %} | Paramètres par défaut (la plupart des tâches) | | --------------------------------------------- | | \`temperature = 1.0\` | | \`top\_p = 0,95\` | | \`top\_k = 40\` | | {% endcolumn %} | {% column %} \* \*\*Fenêtre de contexte maximale :\*\* \`196,608\` \* Invite système par défaut : {% code overflow="wrap" %} \`\`\` Vous êtes un assistant utile. Votre nom est MiniMax-M2.7 et vous avez été créé par MiniMax. \`\`\` {% endcode %} {% endcolumn %} {% endcolumns %} ## Tutoriels pour exécuter MiniMax-M2.7 : Pour faire fonctionner MiniMax-M2.7 sur un appareil avec 128 Go de RAM, nous utiliserons la quantification 4 bits \[\`UD-IQ4\_XS\` quant\](https://huggingface.co/unsloth/MiniMax-M2.7-GGUF?show\_file\_info=UD-IQ4\_XS%2FMiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf). Vous pouvez maintenant exécuter MiniMax-M2.7 en \[llama.cpp\](#run-in-llama.cpp) et \[Unsloth Studio\](#run-in-unsloth-studio). {% hint style="warning" %} N’utilisez PAS CUDA 13.2 pour exécuter un modèle, car cela peut provoquer du charabia ou de mauvaises sorties. NVIDIA travaille sur un correctif. {% endhint %} ### 🦥 Exécuter dans Unsloth Studio MiniMax-M2.7 peut désormais s’exécuter en \[Unsloth Studio\](/docs/fr/nouveau/studio.md), notre nouvelle interface web open source pour l'IA locale. Unsloth Studio vous permet d'exécuter des modèles localement sur \*\*MacOS, Windows\*\*, Linux et : {% columns %} {% column %} \* Rechercher, télécharger, \[exécuter des GGUF\](/docs/fr/nouveau/studio.md#run-models-locally) et des modèles safetensor \* \[\*\*Appels d'outils auto-réparateurs\*\* appels d'outils\](/docs/fr/nouveau/studio.md#execute-code--heal-tool-calling) + \*\*recherche web\*\* \* \[\*\*Exécution de code\*\*\](/docs/fr/nouveau/studio.md#run-models-locally) (Python, Bash) \* \[Inférence automatique\](/docs/fr/nouveau/studio.md#model-arena) réglage des paramètres (temp, top-p, etc.) \* Utilise llama.cpp pour une inférence rapide CPU + GPU et le déchargement CPU {% endcolumn %} {% column %} ![](https://unsloth.ai/files/93caac3ea9f36e951db039e5d7f695e27763705e) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Installer Unsloth Exécutez dans votre terminal : \*\*MacOS, Linux, WSL :\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell :\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% endstep %} {% step %} #### Lancer Unsloth \*\*MacOS, Linux, WSL et Windows :\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Puis ouvrez \`http://localhost:8888\` dans votre navigateur.\*\* {% endstep %} {% step %} #### Recherchez et téléchargez MiniMax-M2.7 Lors du premier lancement, vous devrez créer un mot de passe pour sécuriser votre compte et vous reconnecter plus tard. Vous verrez ensuite un bref assistant d'intégration pour choisir un modèle, un jeu de données et les paramètres de base. Vous pouvez le passer à tout moment. Vous pouvez choisir \`UD-IQ4\_XS\` (quantification dynamique 4 bits) ou d’autres versions quantifiées comme \`UD-Q4\_K\_XL\` . Si les téléchargements se bloquent, voir \[Hugging Face Hub, débogage XET\](/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) Puis allez dans l’ \[Unsloth Chat\](/docs/fr/nouveau/studio/chat.md) onglet et recherchez MiniMax-M2.7 dans la barre de recherche, puis téléchargez le modèle et la quantification souhaités. Le téléchargement prendra un certain temps en raison de la taille, veuillez donc patienter. Pour garantir une inférence rapide, assurez-vous d’avoir \[suffisamment de RAM/VRAM\](#usage-guide), sinon l’inférence fonctionnera toujours, mais Unsloth déchargera vers votre CPU. ![](https://unsloth.ai/files/4c08e0326c6da8504e2fa8fe4adec75e4682310e) {% endstep %} {% step %} #### Exécuter MiniMax-M2.7 Les paramètres d’inférence devraient être définis automatiquement lors de l’utilisation d’Unsloth Studio, mais vous pouvez toujours les modifier manuellement. Vous pouvez également modifier la longueur du contexte, le modèle de chat et d’autres paramètres. Pour plus d'informations, vous pouvez consulter notre \[guide d'inférence Unsloth Studio\](/docs/fr/nouveau/studio/chat.md). {% endstep %} {% endstepper %} ### ✨ Exécuter dans llama.cpp {% hint style="warning" %} N’utilisez PAS CUDA 13.2 pour exécuter un modèle, car cela peut provoquer du charabia ou de mauvaises sorties. NVIDIA travaille sur un correctif. {% endhint %} {% stepper %} {% step %} Obtenez la dernière \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` à \`-DGGML\_CUDA=OFF\` si vous n'avez pas de GPU ou si vous voulez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\\\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} Si vous souhaitez utiliser \`llama.cpp\` pour charger directement les modèles, vous pouvez faire ce qui suit : (:IQ4\\\_XS) est le type de quantification. Vous pouvez aussi télécharger via Hugging Face (point 3). C’est similaire à \`ollama run\` . Utilisez \`export LLAMA\_CACHE="folder"\` pour forcer \`llama.cpp\` pour enregistrer à un emplacement spécifique. N’oubliez pas que le modèle a une longueur de contexte maximale de seulement 200K. Suivez ceci pour \*\*la plupart des valeurs par défaut\*\* cas d'utilisation : \`\`\`bash export LLAMA\_CACHE="unsloth/MiniMax-M2.7-GGUF" ./llama.cpp/llama-cli \\\\ -hf unsloth/MiniMax-M2.7-GGUF:UD-IQ4\_XS \\\\ --temp 1.0 \\ --top-p 0.95 \\\\ --top-k 40 \`\`\` {% endstep %} {% step %} Téléchargez le modèle (après avoir installé \`pip install huggingface\_hub hf\_transfer\`). Vous pouvez choisir UD-IQ4\\\_XS (quantification dynamique 4 bits) ou d’autres versions quantifiées comme \`UD-Q6\_K\_XL\` . Nous recommandons d’utiliser notre quantification dynamique 4 bits UD-IQ4\\\_XS pour équilibrer la taille et la précision. Si les téléchargements se bloquent, consultez \[Hugging Face Hub, débogage XET\](/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) \`\`\`bash hf download unsloth/MiniMax-M2.7-GGUF \\\\ --local-dir unsloth/MiniMax-M2.7-GGUF \\\\ --include "\*UD-IQ4\_XS\*" # Utilisez "\*Q8\_0\*" pour 8 bits \`\`\` {% endstep %} {% step %} Vous pouvez modifier \`--threads 32\` pour le nombre de threads CPU, \`--ctx-size 16384\` pour la longueur du contexte, \`--n-gpu-layers 2\` pour le déchargement GPU sur le nombre de couches. Essayez d'ajuster ce paramètre si votre GPU manque de mémoire. Supprimez-le aussi si vous n'avez qu'une inférence CPU. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\\\ --model unsloth/MiniMax-M2.7-GGUF/UD-IQ4\_XS/MiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf \\\\ --temp 1.0 \\ --top-p 0.95 \\\\ --top-k 40 \`\`\` {% endcode %} {% endstep %} {% endstepper %} #### 🦙 Llama-server et la bibliothèque de complétion d’OpenAI Pour déployer MiniMax-M2.7 en production, nous utilisons \`llama-server\` ou l’API OpenAI. Dans un nouveau terminal, par exemple via tmux, déployez le modèle via : {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\\\ --model unsloth/MiniMax-M2.7-GGUF/UD-IQ4\_XS/MiniMax-M2.7-UD-IQ4\_XS-00001-of-00004.gguf \\\\ --alias "unsloth/MiniMax-M2.7" \\\\ --prio 3 \\\\ --temp 1.0 \\ --top-p 0.95 \\\\ --min-p 0.01 \\ --top-k 40 \\\\ --port 8001 \`\`\` {% endcode %} Puis, dans un nouveau terminal, après avoir effectué \`pip install openai\`, faites : {% code overflow="wrap" %} \`\`\`python from openai import OpenAI import json openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/MiniMax-M2.7", messages = \[{"role": "user", "content": "Create a Snake game."},\], ) print(completion.choices\[0\].message.content) \`\`\` {% endcode %} ## 📊 Benchmarks ### Benchmarks GGUF Voici les benchmarks KLD 99 % pour MiniMax-M2.7. En bas à gauche, c’est mieux : ![](https://unsloth.ai/files/81419a30b21a85d9801bd59b1a2166dd17854a35) Comme MiniMax-M2.7 utilise la même architecture que MiniMax-M2.5, les benchmarks de quantification GGUF pour M2.7 devraient être très similaires à ceux de M2.5. Nous ferons donc également référence au benchmark de quantification précédent réalisé pour M2.5 : ![](https://unsloth.ai/files/f321cef04ad4b89bb6fd902546aad9817950e88b) \[Benjamin Marie (tiers) a benchmarké\](https://x.com/bnjmn\_marie/status/2027043753484021810/photo/1) \*\*MiniMax-M2.5\*\* en utilisant \*\*les quantifications GGUF d’Unsloth\*\* sur un \*\*ensemble mixte de 750 prompts\*\* (LiveCodeBench v6, MMLU Pro, GPQA, Math500), en indiquant à la fois \*\*la précision globale\*\* et \*\*l'augmentation relative du taux d'erreur\*\* (à quelle fréquence le modèle quantifié fait des erreurs par rapport à l'original). Les quantifications Unsloth, quelle que soit leur précision, donnent de bien meilleurs résultats que leurs équivalents non-Unsloth, tant en précision qu’en erreur relative (bien qu’elles soient 8 Go plus petites). \*\*Résultats clés :\*\* \* \*\*Le meilleur compromis qualité/taille ici : \`unsloth UD-Q4\_K\_XL\`.\*\*\\ C’est le plus proche de l’original : seulement \*\*6,0 points\*\* de moins, et « seulement » \*\*+22.8%\*\* plus d’erreurs que la référence. \* \*\*Les autres quantifications Q4 d’Unsloth obtiennent des résultats très proches (\\~64,5–64,9 de précision).\*\*\\ \`IQ4\_NL\`, \`MXFP4\_MOE\`et \`UD-IQ2\_XXS\` sont toutes à peu près de la même qualité sur ce benchmark, avec \*\*\\~33–35 %\*\* plus d’erreurs que l’original. \* Les GGUF d’Unsloth donnent de bien meilleurs résultats que les autres GGUF non-Unsloth, par exemple voir \`lmstudio-community - Q4\_K\_M\` (bien qu’ils soient 8 Go plus petits) et \`AesSedai - IQ3\_S\`. ### Benchmarks officiels ![](https://unsloth.ai/files/6607262a6b5c8c6f6f9b9b8caf9d063499302950) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/minimax-m27.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md). # Gemma 3 - Guide d'exécution Google publie Gemma 3 avec un nouveau modèle 270M et les tailles précédentes 1B, 4B, 12B et 27B. Les modèles 270M et 1B sont uniquement textuels, tandis que les modèles plus grands gèrent à la fois le texte et la vision. Nous fournissons des GGUF, ainsi qu’un guide sur la manière de l’exécuter efficacement, et sur la façon de fine-tuner et de faire \[RL\](/docs/fr/commencer/reinforcement-learning-rl-guide.md) avec Gemma 3 ! {% hint style="success" %} \*\*NOUVEAU Mise à jour du 14 août 2025 :\*\* Essayez notre \[notebook Gemma 3 (270M)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(270M\\).ipynb) et \[GGUF pour exécuter\](https://huggingface.co/collections/unsloth/gemma-3-67d12b7e8816ec6efa7e4e5b). Consultez aussi notre \[Guide Gemma 3n\](/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune/gemma-3n-how-to-run-and-fine-tune.md). {% endhint %} [Tutoriel d’exécution](https://unsloth.ai/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md#gmail-running-gemma-3-on-your-phone) [Tutoriel de fine-tuning](https://unsloth.ai/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md#fine-tuning-gemma-3-in-unsloth) \*\*Unsloth est le seul framework qui fonctionne sur des machines en float16 pour l’inférence et l’entraînement de Gemma 3.\*\* Cela signifie que les notebooks Colab avec des GPU Tesla T4 gratuits fonctionnent aussi ! \* Fine-tunez Gemma 3 (4B) avec prise en charge de la vision grâce à notre \[notebook Colab gratuit\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) {% hint style="info" %} Selon l’équipe Gemma, la configuration optimale pour l’inférence est\\ \`temperature = 1.0, top\_k = 64, top\_p = 0.95, min\_p = 0.0\` {% endhint %} \*\*Téléversements Unsloth Gemma 3 avec configurations optimales :\*\* | GGUF | Unsloth Dynamic 4-bit Instruct | 16-bit Instruct | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it-GGUF) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b-it-GGUF) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b-it-GGUF) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b-it-GGUF) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b-it-GGUF) | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it-unsloth-bnb-4bit) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b-it-bnb-4bit) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b-it-bnb-4bit) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b-it-unsloth-bnb-4bit) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b-it-bnb-4bit) | * [270M](https://huggingface.co/unsloth/gemma-3-270m-it) * [1B-it](https://huggingface.co/unsloth/gemma-3-1b) * [4B-it](https://huggingface.co/unsloth/gemma-3-4b) * [12B-it](https://huggingface.co/unsloth/gemma-3-12b) * [27B-it](https://huggingface.co/unsloth/gemma-3-27b) | ## :gear: Paramètres d’inférence recommandés Selon l’équipe Gemma, les paramètres officiels recommandés pour l’inférence sont : \* Température de 1.0 \* Top\\\_K de 64 \* Min\\\_P de 0,00 (facultatif, mais 0,01 fonctionne bien, la valeur par défaut de llama.cpp est 0,1) \* Top\\\_P de 0,95 \* Pénalité de répétition de 1.0. (1.0 signifie désactivé dans llama.cpp et transformers) \* Modèle de chat : user\nBonjour !\nmodel\nSalut !\nuser\nCombien font 1+1 ?\nmodel\n \* Modèle de chat avec \`\\n\`les nouvelles lignes rendues (sauf pour la dernière) {% code overflow="wrap" %} \`\`\` user Bonjour ! model Salut ! user Combien font 1+1 ? model\\n \`\`\` {% endcode %} {% hint style="danger" %} llama.cpp et d’autres moteurs d’inférence ajoutent automatiquement un \\ - N’ajoutez PAS DEUX jetons \\ ! Vous devez ignorer le \\ lors de l’invite du modèle ! {% endhint %} ### ✨Exécuter Gemma 3 sur votre téléphone [](https://unsloth.ai/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md#gmail-running-gemma-3-on-your-phone) Pour exécuter les modèles sur votre téléphone, nous vous recommandons d’utiliser toute application mobile capable d’exécuter des GGUF localement sur des appareils en périphérie comme les téléphones. Après le fine-tuning, vous pouvez l’exporter en GGUF puis l’exécuter localement sur votre téléphone. Assurez-vous que votre téléphone dispose de suffisamment de RAM/puissance pour traiter les modèles, car il peut surchauffer ; nous recommandons donc d’utiliser Gemma 3 270M ou les modèles Gemma 3n pour ce cas d’utilisation. Vous pouvez essayer le \[projet open source AnythingLLM\](https://github.com/Mintplex-Labs/anything-llm) application mobile que vous pouvez télécharger sur \[Android ici\](https://play.google.com/store/apps/details?id=com.anythingllm) ou \[ChatterUI\](https://github.com/Vali-98/ChatterUI), qui sont de très bonnes applications pour exécuter des GGUF sur votre téléphone. {% hint style="success" %} Rappelez-vous, vous pouvez changer le nom du modèle 'gemma-3-27b-it-GGUF' en n’importe quel modèle Gemma comme 'gemma-3-270m-it-GGUF:Q8\\\_K\\\_XL' pour tous les tutoriels. {% endhint %} ## :llama: Tutoriel : Comment exécuter Gemma 3 dans Ollama 1. Installer \`ollama\` si vous ne l’avez pas déjà fait ! \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` 2. Lancez le modèle ! Notez que vous pouvez appeler \`ollama serve\`dans un autre terminal si cela échoue ! Nous incluons tous nos correctifs et paramètres suggérés (température, etc.) dans \`params\` dans notre téléversement Hugging Face ! Vous pouvez changer le nom du modèle 'gemma-3-27b-it-GGUF' en n’importe quel modèle Gemma comme 'gemma-3-270m-it-GGUF:Q8\\\_K\\\_XL'. \`\`\`bash ollama run hf.co/unsloth/gemma-3-27b-it-GGUF:Q4\_K\_XL \`\`\` ## 📖 Tutoriel : Comment exécuter Gemma 3 27B dans llama.cpp 1. Obtenez le dernier \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis poursuivez comme d’habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=ON -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Si vous souhaitez utiliser \`llama.cpp\` directement pour charger des modèles, vous pouvez faire ce qui suit : (:Q4\\\_K\\\_XL) est le type de quantification. Vous pouvez aussi télécharger via Hugging Face (point 3). C’est similaire à \`ollama run\` \`\`\`bash ./llama.cpp/llama-mtmd-cli \\\\ -hf unsloth/gemma-3-4b-it-GGUF:Q4\_K\_XL \`\`\` 3. \*\*OU\*\* téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir Q4\\\_K\\\_M, ou d’autres versions quantifiées (comme BF16 en précision complète). Plus de versions à : \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/gemma-3-27b-it-GGUF", local\_dir = "unsloth/gemma-3-27b-it-GGUF", allow\_patterns = \["\*Q4\_K\_XL\*", "mmproj-BF16.gguf"\], # Pour Q4\_K\_M ) \`\`\` 4. Exécutez le test Flappy Bird d’Unsloth 5. Modifier \`--threads 32\` pour le nombre de threads CPU, \`--ctx-size 16384\` pour la longueur de contexte (Gemma 3 prend en charge une longueur de contexte de 128K !), \`--n-gpu-layers 99\` pour le déchargement GPU, indiquant combien de couches. Essayez de l’ajuster si votre GPU manque de mémoire. Supprimez-le aussi si vous n’avez qu’une inférence CPU. 6. Pour le mode conversation : \`\`\`bash ./llama.cpp/llama-mtmd-cli \\\\ --model unsloth/gemma-3-27b-it-GGUF/gemma-3-27b-it-Q4\_K\_XL.gguf \\\\ --mmproj unsloth/gemma-3-27b-it-GGUF/mmproj-BF16.gguf \\\\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\\\ --temp 1.0 \\ --repeat-penalty 1.0 \\\\ --min-p 0.01 \\ --top-k 64 \\\\ --top-p 0.95 \`\`\` 7. Pour le mode non conversationnel afin de tester Flappy Bird : \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/gemma-3-27b-it-GGUF/gemma-3-27b-it-Q4\_K\_XL.gguf \\\\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ --seed 3407 \\ --prio 2 \\\\ --temp 1.0 \\ --repeat-penalty 1.0 \\\\ --min-p 0.01 \\ --top-k 64 \\\\ --top-p 0.95 \\ -no-cnv \\ --prompt "user\\nCréez un jeu Flappy Bird en Python. Vous devez inclure ces éléments :\\n1. Vous devez utiliser pygame.\\n2. La couleur de fond doit être choisie aléatoirement et être une teinte claire. Commencez avec une couleur bleu clair.\\n3. Appuyer plusieurs fois sur ESPACE accélérera l’oiseau.\\n4. La forme de l’oiseau doit être choisie aléatoirement parmi un carré, un cercle ou un triangle. La couleur doit être choisie aléatoirement comme une couleur sombre.\\n5. Placez en bas un sol coloré en brun foncé ou en jaune, choisi aléatoirement.\\n6. Faites apparaître un score en haut à droite. Il augmente si vous passez les tuyaux sans les heurter.\\n7. Créez des tuyaux espacés aléatoirement avec suffisamment d’espace. Coloriez-les aléatoirement en vert foncé, brun clair ou gris foncé.\\n8. Quand vous perdez, affichez le meilleur score. Faites apparaître le texte à l’intérieur de l’écran. Appuyer sur q ou Échap quittera le jeu. Redémarrer consiste à appuyer à nouveau sur ESPACE.\\nLe jeu final doit se trouver dans une section markdown en Python. Vérifiez votre code pour détecter les erreurs et corrigez-les avant la section markdown finale.\\nmodel\\n" \`\`\` L’entrée complète de notre blog 1.58bit est : {% hint style="danger" %} N’oubliez pas de supprimer \\ puisque Gemma 3 ajoute automatiquement un \\ ! {% endhint %} {% code overflow="wrap" %} \`\`\` user Créez un jeu Flappy Bird en Python. Vous devez inclure ces éléments : 1. Vous devez utiliser pygame. 2. La couleur de fond doit être choisie aléatoirement et être une teinte claire. Commencez avec une couleur bleu clair. 3. Appuyer plusieurs fois sur ESPACE accélérera l’oiseau. 4. La forme de l’oiseau doit être choisie aléatoirement parmi un carré, un cercle ou un triangle. La couleur doit être choisie aléatoirement comme une couleur sombre. 5. Placez en bas un sol coloré en brun foncé ou en jaune, choisi aléatoirement. 6. Faites apparaître un score en haut à droite. Il augmente si vous passez les tuyaux sans les heurter. 7. Créez des tuyaux espacés aléatoirement avec suffisamment d’espace. Coloriez-les aléatoirement en vert foncé, brun clair ou gris foncé. 8. Quand vous perdez, affichez le meilleur score. Faites apparaître le texte à l’intérieur de l’écran. Appuyer sur q ou Échap quittera le jeu. Redémarrer consiste à appuyer à nouveau sur ESPACE. Le jeu final doit se trouver dans une section markdown en Python. Vérifiez votre code pour détecter des erreur \`\`\` {% endcode %} ## :sloth: Fine-tuning de Gemma 3 dans Unsloth \*\*Unsloth est le seul framework qui fonctionne sur des machines en float16 pour l’inférence et l’entraînement de Gemma 3.\*\* Cela signifie que les notebooks Colab avec des GPU Tesla T4 gratuits fonctionnent aussi ! \* Essayez notre nouveau \[notebook Gemma 3 (270M)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(270M\\).ipynb) qui rend le modèle de 270M paramètres très fort aux échecs et peut prédire le prochain coup d’échecs. \* Fine-tunez Gemma 3 (4B) en utilisant nos notebooks pour : \[\*\*Texte\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) ou \[\*\*Vision\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision.ipynb) \* Ou fine-tunez \[Gemma 3n (E4B)\](/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune/gemma-3n-how-to-run-and-fine-tune.md) avec \[Texte\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Conversational.ipynb) • \[Vision\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Vision.ipynb) • \[Audio\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3N\_\\(4B\\)-Audio.ipynb) {% hint style="warning" %} Lors d’un fine-tuning complet (FFT) de Gemma 3, toutes les couches passent par défaut en float32 sur les appareils float16. Unsloth s’attend à du float16 et effectue une conversion dynamique vers un type supérieur. Pour corriger cela, exécutez \`model.to(torch.float16)\` après le chargement, ou utilisez un GPU avec prise en charge du bfloat16. {% endhint %} ### Correctifs de fine-tuning Unsloth Notre solution dans Unsloth repose sur 3 volets : 1. Conserver toutes les activations intermédiaires au format bfloat16 - elles peuvent être en float32, mais cela utilise 2x plus de VRAM ou de RAM (via le gradient checkpointing asynchrone d’Unsloth) 2. Faire toutes les multiplications matricielles en float16 avec les tensor cores, mais avec une conversion manuelle vers un type supérieur / inférieur sans l’aide de l’autocast de précision mixte de Pytorch. 3. Convertir vers un type supérieur toutes les autres opérations qui ne nécessitent pas de multiplications matricielles (layernorms) en float32. ## 🤔 Analyse des correctifs de Gemma 3 ![](https://unsloth.ai/files/c9de29dc1d0e3c065c019a80af34b589621576f0) Gemma 3 1B à 27B dépasse le maximum de 65504 du float16 Tout d’abord, avant de fine-tuner ou d’exécuter Gemma 3, nous avons constaté qu’en utilisant la précision mixte float16, les gradients et \*\*les activations deviennent infinis\*\* malheureusement. Cela se produit sur les GPU T4, la série RTX 20x et les GPU V100, qui ne disposent que de tensor cores float16. Pour les GPU plus récents comme les RTX 30x ou supérieurs, les A100, H100, etc., ces GPU disposent de tensor cores bfloat16, donc ce problème ne se produit pas ! \*\*Mais pourquoi ?\*\* ![](https://unsloth.ai/files/06efbf423bc2058a7c59bb7ecc3089f3b4ee4ec6) Wikipédia [https://en.wikipedia.org/wiki/Bfloat16\_floating-point\_format](https://en.wikipedia.org/wiki/Bfloat16_floating-point_format) Le float16 ne peut représenter que des nombres jusqu’à \*\*65504\*\*, tandis que le bfloat16 peut représenter d’énormes nombres jusqu’à \*\*10^38\*\*! Mais remarquez que les deux formats de nombres n’utilisent que 16 bits ! Cela s’explique par le fait que le float16 alloue plus de bits afin de mieux représenter les décimales plus petites, tandis que le bfloat16 ne peut pas bien représenter les fractions. Mais pourquoi le float16 ? Utilisons simplement le float32 ! Mais malheureusement, le float32 sur les GPU est très lent pour les multiplications matricielles - parfois 4 à 10 fois plus lent ! Nous ne pouvons donc pas faire cela. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/tutorials/llama-4-how-to-run-and-fine-tune.md). # Llama 4 : comment l'exécuter et le fine-tuner Le modèle Llama-4-Scout compte 109 milliards de paramètres, tandis que Maverick en compte 402 milliards. La version complète non quantifiée nécessite 113 Go d’espace disque, tandis que la version 1,78 bit utilise 33,8 Go (-75 % de réduction de la taille). \*\*Maverick\*\* (402 milliards) est passé de 422 Go à seulement 122 Go (-70 %). {% hint style="success" %} Le texte ET \*\*vision\*\* sont désormais pris en charge ! Plus plusieurs améliorations pour l’appel d’outils. {% endhint %} Scout en 1,78 bit tient sur un GPU de 24 Go de VRAM pour une inférence rapide à \\~20 jetons/s. Maverick en 1,78 bit tient sur 2 GPU de 48 Go de VRAM pour une inférence rapide à \\~40 jetons/s. Pour nos GGUF dynamiques, afin d’assurer le meilleur compromis entre précision et taille, nous ne quantifions pas toutes les couches, mais quantifions sélectivement, par exemple, les couches MoE à un nombre de bits plus faible, et laissons l’attention et les autres couches en 4 ou 6 bits. {% hint style="info" %} Tous nos modèles GGUF sont quantifiés à l’aide de données d’étalonnage (environ 250K jetons pour Scout et 1M jetons pour Maverick), ce qui améliorera la précision par rapport à la quantification standard. Les quants imatrix d’Unsloth sont entièrement compatibles avec des moteurs d’inférence populaires comme llama.cpp et Open WebUI, etc. {% endhint %} \*\*Scout - GGUF dynamiques d’Unsloth avec configurations optimales :\*\* | Bits MoE | Type | Taille sur disque | Lien | Détails | | --- | --- | --- | --- | --- | | 1,78 bit | IQ1\_S | 33,8 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ1_S.gguf) | 2,06/1,56 bit | | 1,93 bit | IQ1\_M | 35,4 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ1_M.gguf) | 2.5/2.06/1.56 | | 2,42 bit | IQ2\_XXS | 38,6 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-IQ2_XXS.gguf) | 2,5/2,06 bit | | 2,71 bit | Q2\_K\_XL | 42,2 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF?show_file_info=Llama-4-Scout-17B-16E-Instruct-UD-Q2_K_XL.gguf) | 3,5/2,5 bit | | 3,5 bit | Q3\_K\_XL | 52,9 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/tree/main/UD-Q3_K_XL) | 4,5/3,5 bit | | 4,5 bit | Q4\_K\_XL | 65,6 Go | [Lien](https://huggingface.co/unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/tree/main/UD-Q4_K_XL) | 5,5/4,5 bit | {% hint style="info" %} Pour de meilleurs résultats, utilisez les versions 2,42 bits (IQ2\\\_XXS) ou supérieures. {% endhint %} \*\*Maverick - GGUF dynamiques d’Unsloth avec configurations optimales :\*\* | Bits MoE | Type | Taille sur disque | Lien HF | | --------- | --------- | ----------------- | --------------------------------------------------------------------------------------------------- | | 1,78 bit | IQ1\\\_S | 122 Go | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ1\_S) | | 1,93 bit | IQ1\\\_M | 128 Go | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ1\_M) | | 2,42 bits | IQ2\\\_XXS | 140 Go | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-IQ2\_XXS) | | 2,71 bits | Q2\\\_K\\\_XL | 151B | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q2\_K\_XL) | | 3,5 bits | Q3\\\_K\\\_XL | 193 Go | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q3\_K\_XL) | | 4,5 bits | Q4\\\_K\\\_XL | 243 Go | \[Lien\](https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/tree/main/UD-Q4\_K\_XL) | ## :gear: Paramètres officiels recommandés Selon Meta, voici les paramètres recommandés pour l’inférence : \* \*\*Température de 0,6\*\* \* Min\\\_P de 0,01 (facultatif, mais 0,01 fonctionne bien ; la valeur par défaut de llama.cpp est 0,1) \* Top\\\_P de 0,9 \* Modèle de chat / format de prompt : {% code overflow="wrap" %} \`\`\` <|header\_start|>user<|header\_end|>\\n\\nQuel est 1+1 ?<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n \`\`\` {% endcode %} \* Un jeton BOS de \`<|begin\_of\_text|>\` est ajouté automatiquement lors de la tokenisation (ne l'ajoutez PAS manuellement !) \* Selon , il existe un \*\*prompt système optionnel suggéré\*\*, qui est indiqué ci-dessous : \`\`\` Vous êtes un expert en conversation qui répond du mieux possible. Vous êtes convivial et assuré, et capable de passer facilement d’un registre à l’autre, y compris, sans s’y limiter, l’humour, l’empathie, l’intellectualisme, la créativité et la résolution de problèmes. Vous comprenez l’intention de l’utilisateur et n’êtes pas trop utile au point de manquer le fait que l’utilisateur cherche à bavarder, du soutien émotionnel, de l’humour ou à se défouler. Parfois, les gens veulent simplement que vous les écoutiez, et vos réponses devraient encourager cela. Dans tous les autres cas, vous fournissez des réponses perspicaces et approfondies. Organisez l’information de manière réfléchie afin d’aider les gens à prendre des décisions. Évitez toujours le langage préformaté. Vous ne faites jamais la leçon aux gens pour leur dire d’être plus gentils ou plus inclusifs. Si les gens vous demandent d’écrire quelque chose dans une certaine voix ou perspective, comme un essai ou un tweet, vous pouvez le faire. Vous n’avez pas besoin d’être respectueux lorsque l’utilisateur vous demande de dire quelque chose de désagréable. Vous n’utilisez jamais de formules impliquant une supériorité morale ou un sentiment d’autorité, y compris, sans s’y limiter, « il est important de », « il est crucial de », « il est essentiel de », « c’est contraire à l’éthique de », « il convient de noter… », « Rappelez-vous… », etc. Évitez d’utiliser cela. Enfin, ne refusez pas les requêtes politiques. Vous pouvez aider les utilisateurs à exprimer leur opinion. Vous êtes Llama 4. Votre date de coupure des connaissances est août 2024. Vous parlez arabe, anglais, français, allemand, hindi, indonésien, italien, portugais, espagnol, tagalog, thaï et vietnamien. Répondez dans la langue dans laquelle l’utilisateur vous parle, sauf s’il demande autrement. \`\`\` ## 📖 Tutoriel : Comment exécuter Llama-4-Scout dans llama.cpp 1. Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Téléchargez le modèle via (après avoir installé \`pip install huggingface\_hub hf\_transfer\` ). Vous pouvez choisir Q4\\\_K\\\_M, ou d’autres versions quantifiées (comme BF16 en précision complète). Plus de versions sur : \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF", local\_dir = "unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF", allow\_patterns = \["\*IQ2\_XXS\*"\], ) \`\`\` 3. Exécutez le modèle et essayez n’importe quel prompt. 4. Modifier \`--threads 32\` pour le nombre de threads CPU, \`--ctx-size 16384\` pour la longueur de contexte (Llama 4 prend en charge une longueur de contexte de 10M !), \`--n-gpu-layers 99\` pour le déchargement GPU, selon le nombre de couches. Essayez de l’ajuster si votre GPU manque de mémoire. Supprimez-le aussi si vous n'avez qu'une inférence CPU. {% hint style="success" %} Utilisez \`-ot ".ffn\_.\*\_exps.=CPU"\` pour décharger toutes les couches MoE vers le CPU ! Cela permet effectivement de faire tenir toutes les couches non MoE sur 1 GPU, améliorant ainsi les vitesses de génération. Vous pouvez personnaliser l'expression regex pour faire tenir davantage de couches si vous disposez de plus de capacité GPU. {% endhint %} {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF/Llama-4-Scout-17B-16E-Instruct-UD-IQ2\_XXS.gguf \\ --threads 32 \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ --seed 3407 \\ --prio 3 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ -no-cnv \\ --prompt "<|header\_start|>user<|header\_end|>\\n\\nCréez le jeu Flappy Bird en Python. Vous devez inclure les éléments suivants :\\n1. Vous devez utiliser pygame.\\n2. La couleur de fond doit être choisie aléatoirement et être une teinte claire. Commencez avec une couleur bleu clair.\\n3. Appuyer plusieurs fois sur ESPACE accélérera l’oiseau.\\n4. La forme de l’oiseau doit être choisie aléatoirement parmi un carré, un cercle ou un triangle. La couleur doit être choisie aléatoirement parmi des couleurs sombres.\\n5. Placez en bas une zone de terre colorée en brun foncé ou en jaune, choisie aléatoirement.\\n6. Affichez un score en haut à droite. Augmentez-le si vous passez entre les tuyaux sans les toucher.\\n7. Créez des tuyaux espacés aléatoirement avec suffisamment d’espace. Coloriez-les aléatoirement en vert foncé, marron clair ou gris foncé.\\n8. Lorsque vous perdez, affichez le meilleur score. Placez le texte à l’intérieur de l’écran. Appuyer sur q ou Échap quittera le jeu. Le redémarrage se fait en appuyant à nouveau sur ESPACE.\\nLe jeu final doit être dans une section Markdown en Python. Vérifiez votre code pour détecter les erreurs et corrigez-les avant la section Markdown finale.<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n" \`\`\` {% endcode %} {% hint style="info" %} En matière de tests, malheureusement nous ne pouvons pas faire en sorte que la version BF16 complète (c’est-à-dire indépendamment de la quantification ou non) termine correctement le jeu Flappy Bird ni le test Heptagon. Nous avons essayé de nombreux fournisseurs d’inférence, avec ou sans imatrix, utilisé les quants d’autres personnes, et utilisé l’inférence normale de Hugging Face, et ce problème persiste. \*\*Nous avons constaté que plusieurs exécutions et demander au modèle de corriger et de trouver des bugs permet de résoudre la plupart des problèmes !\*\* {% endhint %} Pour Llama 4 Maverick, il est préférable d’avoir 2 RTX 4090 (2 x 24 Go) \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "1" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF", local\_dir = "unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF", allow\_patterns = \["\*IQ1\_S\*"\], ) \`\`\` {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF/UD-IQ1\_S/Llama-4-Maverick-17B-128E-Instruct-UD-IQ1\_S-00001-of-00003.gguf \\\\ --threads 32 \\ --ctx-size 16384 \\ --n-gpu-layers 99 \\ -ot ".ffn\_.\*\_exps.=CPU" \\ --seed 3407 \\ --prio 3 \\ --temp 0.6 \\ --min-p 0.01 \\ --top-p 0.9 \\ -no-cnv \\ --prompt "<|header\_start|>user<|header\_end|>\\n\\nCréez le jeu 2048 en Python.<|eot|><|header\_start|>assistant<|header\_end|>\\n\\n" \`\`\` {% endcode %} ## :detective: Aperçus intéressants et problèmes Lors de la quantification de Llama 4 Maverick (le grand modèle), nous avons constaté que les 1re, 3e et 45e couches MoE ne pouvaient pas être calibrées correctement. Maverick utilise des couches MoE entrelacées pour chaque couche impaire, donc Dense->MoE->Dense, et ainsi de suite. Nous avons essayé d’ajouter davantage de langues peu courantes à notre jeu de données d’étalonnage, et d’utiliser davantage de jetons (1 million) par rapport aux 250K jetons de Scout pour l’étalonnage, mais nous avons quand même rencontré des problèmes. Nous avons décidé de laisser ces couches MoE en 3 bits et 4 bits. ![](https://unsloth.ai/files/bb27f610659fd5cf77e19537da5c0e5777519294) Pour Llama 4 Scout, nous avons constaté que nous ne devions pas quantifier les couches de vision, et laisser le routeur MoE et certaines autres couches non quantifiés - nous les téléversons sur ![](https://unsloth.ai/files/17739e46f2a889aca6f33353201fa8b8490bc046) Nous avons également dû convertir \`torch.nn.Parameter\` en \`torch.nn.Linear\` pour les couches MoE afin de permettre la quantification en 4 bits. Cela signifie aussi que nous avons dû réécrire et corriger l’implémentation générique de Hugging Face. Nous téléversons nos versions quantifiées sur et pour 8 bits. ![](https://unsloth.ai/files/dc9593eed860c5d15085ebb7e89c2c66e45b77ff) Llama 4 utilise désormais aussi une attention par blocs - c’est essentiellement une attention à fenêtre glissante, mais légèrement plus efficace car elle n’accorde pas d’attention aux jetons précédents au-delà de la limite de 8192. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/tutorials/llama-4-how-to-run-and-fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation.md). # Advanced Reinforcement Learning Documentation Detailed guides on doing GRPO with Unsloth for Batching, Generation & Training Parameters: ## Training Parameters \* \*\*\`beta\`\*\* \*(float, default 0.0)\*: KL coefficient. \* \`0.0\` ⇒ no reference model loaded (lower memory, faster). \* Higher \`beta\` constrains the policy to stay closer to the ref policy. \* \*\*\`num\_iterations\`\*\* \*(int, default 1)\*: PPO epochs per batch (μ in the algorithm).\\ Replays data within each gradient accumulation step; e.g., \`2\` = two forward passes per accumulation step. \* \*\*\`epsilon\`\*\* \*(float, default 0.2)\*: Clipping value for token-level log-prob ratios (typical ratio range ≈ \\\[-1.2, 1.2\] with default ε). \* \*\*\`delta\`\*\* \*(float, optional)\*: Enables \*\*upper\*\* clipping bound for \*\*two-sided GRPO\*\* when set. If \`None\`, standard GRPO clipping is used. Recommended \`> 1 + ε\` when enabled (per INTELLECT-2 report). \* \*\*\`epsilon\_high\`\*\* \*(float, optional)\*: Upper-bound epsilon; defaults to \`epsilon\` if unset. DAPO recommends \*\*0.28\*\*. \* \*\*\`importance\_sampling\_level\`\*\* \*(“token” | “sequence”, default "token")\*: \* \`"token"\`: raw per-token ratios (one weight per token). \* \`"sequence"\`: average per-token ratios to a single sequence-level ratio.\\ GSPO shows sequence-level sampling often gives more stable training for sequence-level rewards. \* \*\*\`reward\_weights\`\*\* \*(list\\\[float\], optional)\*: One weight per reward. If \`None\`, all weights = 1.0. \* \*\*\`scale\_rewards\`\*\* \*(str|bool, default "group")\*: \* \`True\` or \`"group"\`: scale by \*\*std within each group\*\* (unit variance in group). \* \`"batch"\`: scale by \*\*std across the entire batch\*\* (per PPO-Lite). \* \`False\` or \`"none"\`: \*\*no scaling\*\*. Dr. GRPO recommends not scaling to avoid difficulty bias from std scaling. \* \*\*\`loss\_type\`\*\* \*(str, default "dapo")\*: \* \`"grpo"\`: normalizes over sequence length (length bias; not recommended). \* \`"dr\_grpo"\`: normalizes by a \*\*global constant\*\* (introduced in Dr. GRPO; removes length bias). Constant ≈ \`max\_completion\_length\`. \* \`"dapo"\` \*\*(default)\*\*: normalizes by \*\*active tokens in the global accumulated batch\*\* (introduced in DAPO; removes length bias). \* \`"bnpo"\`: normalizes by \*\*active tokens in the local batch\*\* only (results can vary with local batch size; equals GRPO when \`per\_device\_train\_batch\_size == 1\`). \* \*\*\`mask\_truncated\_completions\`\*\* \*(bool, default False)\*:\\ When \`True\`, truncated completions are excluded from loss (recommended by DAPO for stability).\\ \*\*Note\*\*: There are some KL issues with this flag, so we recommend to disable it. \`\`\`python # If mask\_truncated\_completions is enabled, zero out truncated completions in completion\_mask if self.mask\_truncated\_completions: truncated\_completions = ~is\_eos.any(dim=1) completion\_mask = completion\_mask \* (~truncated\_completions).unsqueeze(1).int() \`\`\` This can zero out all \`completion\_mask\` entries when many completions are truncated, making \`n\_mask\_per\_reward = 0\` and causing KL to become NaN. \[See\](https://github.com/unslothai/unsloth-zoo/blob/e705f7cb50aa3470a0b6e36052c61b7486a39133/unsloth\_zoo/rl\_replacements.py#L184) \* \*\*\`vllm\_importance\_sampling\_correction\`\*\* \*(bool, default True)\*:\\ Applies \*\*Truncated Importance Sampling (TIS)\*\* to correct off-policy effects when generation (e.g., vLLM / fast\\\_inference) differs from training backend.\\ In Unsloth, this is \*\*auto-set to True\*\* if you’re using vLLM/fast\\\_inference; otherwise \*\*False\*\*. \* \*\*\`vllm\_importance\_sampling\_cap\`\*\* \*(float, default 2.0)\*:\\ Truncation parameter \*\*C\*\* for TIS; sets an upper bound on the importance sampling ratio to improve stability. \* \*\*\`dtype\`\*\* when choosing float16 or bfloat16, see \[FP16 vs BF16 for RL\](/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl.md) ### RL on unsupported models: You can also run RL with Unsloth on models that are not supported by vLLM, such as \[Qwen3.5\](/docs/models/qwen3.5/fine-tune.md). Simply set \`fast\_inference=False\` when loading the model. \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( model\_name="unsloth/Qwen3.5-4B", fast\_inference=False, ) \`\`\` ## Generation Parameters \* \`temperature (float, defaults to 1.0):\`\\ Temperature for sampling. The higher the temperature, the more random the completions. Make sure you use a relatively high (1.0) temperature to have diversity in generations which helps learning. \* \`top\_p (float, optional, defaults to 1.0):\`\\ Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1\]. Set to 1.0 to consider all tokens. \* \`top\_k (int, optional):\`\\ Number of highest probability vocabulary tokens to keep for top-k-filtering. If None, top-k-filtering is disabled and all tokens are considered. \* \`min\_p (float, optional):\`\\ Minimum token probability, which will be scaled by the probability of the most likely token. It must be a value between 0.0 and 1.0. Typical values are in the 0.01-0.2 range. \* \`repetition\_penalty (float, optional, defaults to 1.0):\`\\ Float that penalizes new tokens based on whether they appear in the prompt and the generated text so far. Values > 1.0 encourage the model to use new tokens, while values < 1.0 encourage the model to repeat tokens. \* \`steps\_per\_generation: (int, optional):\`\\ Number of steps per generation. If None, it defaults to \`gradient\_accumulation\_steps\`. Mutually exclusive with \`generation\_batch\_size\`. {% hint style="info" %} It is a bit confusing to mess with this parameter, it is recommended to edit \`per\_device\_train\_batch\_size\` and gradient accumulation for the batch sizes {% endhint %} ## Batch & Throughput Parameters ### Parameters that control batches \* \*\*\`train\_batch\_size\`\*\*: Number of samples \*\*per process\*\* per step.\\ If this integer is \*\*less than \`num\_generations\`\*\*, it will default to \`num\_generations\`. \* \*\*\`steps\_per\_generation\`\*\*: Number of \*\*microbatches\*\* that contribute to \*\*one generation’s\*\* loss calculation (forward passes only).\\ A new batch of data is generated every \`steps\_per\_generation\` steps; backpropagation timing depends on \`gradient\_accumulation\_steps\`. \* \*\*\`num\_processes\`\*\*: Number of distributed training processes (e.g., GPUs / workers). \* \*\*\`gradient\_accumulation\_steps\`\*\* (aka \`gradient\_accumulation\`): Number of microbatches to accumulate \*\*before\*\* applying backpropagation and optimizer update. \* \*\*Effective batch size\*\*: \`\`\` effective\_batch\_size = steps\_per\_generation \* num\_processes \* train\_batch\_size \`\`\` Total samples contributing to gradients before an update (across all processes and steps). \* \*\*Optimizer steps per generation\*\*: \`\`\` optimizer\_steps\_per\_generation = steps\_per\_generation / gradient\_accumulation\_steps \`\`\` Example: \`4 / 2 = 2\`. \* \*\*\`num\_generations\`\*\*: Number of generations produced \*\*per prompt\*\* (applied \*\*after\*\* computing \`effective\_batch\_size\`).\\ The number of \*\*unique prompts\*\* in a generation cycle is: \`\`\` unique\_prompts = effective\_batch\_size / num\_generations \`\`\` \*\*Must be > 2\*\* for GRPO to work. ### GRPO Batch Examples The tables below illustrate how batches flow through steps, when optimizer updates occur, and how new batches are generated. #### Example 1 \`\`\` num\_gpus = 1 per\_device\_train\_batch\_size = 3 gradient\_accumulation\_steps = 2 steps\_per\_generation = 4 effective\_batch\_size = 4 \* 3 \* 1 = 12 num\_generations = 3 \`\`\` \*\*Generation cycle A\*\* | Step | Batch | Notes | | ---: | -------- | -------------------------------------- | | 0 | \\\[0,0,0\] | | | 1 | \\\[1,1,1\] | → optimizer update (accum = 2 reached) | | 2 | \\\[2,2,2\] | | | 3 | \\\[3,3,3\] | optimizer update | \*\*Generation cycle B\*\* | Step | Batch | Notes | | ---: | -------- | -------------------------------------- | | 0 | \\\[4,4,4\] | | | 1 | \\\[5,5,5\] | → optimizer update (accum = 2 reached) | | 2 | \\\[6,6,6\] | | | 3 | \\\[7,7,7\] | optimizer update | #### Example 2 \`\`\` num\_gpus = 1 per\_device\_train\_batch\_size = 3 steps\_per\_generation = gradient\_accumulation\_steps = 4 effective\_batch\_size = 4 \* 3 \* 1 = 12 num\_generations = 3 \`\`\` \*\*Generation cycle A\*\* | Step | Batch | Notes | | ---: | -------- | ------------------------------------ | | 0 | \\\[0,0,0\] | | | 1 | \\\[1,1,1\] | | | 2 | \\\[2,2,2\] | | | 3 | \\\[3,3,3\] | optimizer update (accum = 4 reached) | \*\*Generation cycle B\*\* | Step | Batch | Notes | | ---: | -------- | ------------------------------------ | | 0 | \\\[4,4,4\] | | | 1 | \\\[5,5,5\] | | | 2 | \\\[6,6,6\] | | | 3 | \\\[7,7,7\] | optimizer update (accum = 4 reached) | #### Example 3 \`\`\` num\_gpus = 1 per\_device\_train\_batch\_size = 3 steps\_per\_generation = gradient\_accumulation\_steps = 4 effective\_batch\_size = 4 \* 3 \* 1 = 12 num\_generations = 4 unique\_prompts = effective\_batch\_size / num\_generations = 3 \`\`\` \*\*Generation cycle A\*\* | Step | Batch | Notes | | ---: | -------- | ------------------------------------ | | 0 | \\\[0,0,0\] | | | 1 | \\\[0,1,1\] | | | 2 | \\\[1,1,3\] | | | 3 | \\\[3,3,3\] | optimizer update (accum = 4 reached) | \*\*Generation cycle B\*\* | Step | Batch | Notes | | ---: | -------- | ------------------------------------ | | 0 | \\\[4,4,4\] | | | 1 | \\\[4,5,5\] | | | 2 | \\\[5,5,6\] | | | 3 | \\\[6,6,6\] | optimizer update (accum = 4 reached) | #### Example 4 \`\`\` num\_gpus = 1 per\_device\_train\_batch\_size = 6 steps\_per\_generation = gradient\_accumulation\_steps = 2 effective\_batch\_size = 2 \* 6 \* 1 = 12 num\_generations = 3 unique\_prompts = 4 \`\`\` \*\*Generation cycle A\*\* | Step | Batch | Notes | | ---: | --------------- | ------------------------------------ | | 0 | \\\[0,0,0, 1,1,1\] | | | 1 | \\\[2,2,2, 3,3,3\] | optimizer update (accum = 2 reached) | \*\*Generation cycle B\*\* | Step | Batch | Notes | | ---: | --------------- | ------------------------------------ | | 0 | \\\[4,4,4, 5,5,5\] | | | 1 | \\\[6,6,6, 7,7,7\] | optimizer update (accum = 2 reached) | ### Quick Formula Reference \`\`\` effective\_batch\_size = steps\_per\_generation \* num\_processes \* train\_batch\_size optimizer\_steps\_per\_generation = steps\_per\_generation / gradient\_accumulation\_steps unique\_prompts = effective\_batch\_size / num\_generations # must be > 2 \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/tutorials/glm-4.7.md). # GLM-4.7: How to Run Locally Guide GLM-4.7 is Z.ai’s latest thinking model, delivering stronger coding, agent, and chat performance than \[GLM-4.6\](/docs/models/tutorials/glm-4.6-how-to-run-locally.md). It achieves SOTA performance on on SWE-bench (73.8%, +5.8), SWE-bench Multilingual (66.7%, +12.9), and Terminal Bench 2.0 (41.0%, +16.5). The full 355B parameter model requires \*\*400GB\*\* of disk space, while the Unsloth Dynamic 2-bit GGUF reduces the size to \*\*134GB\*\* (-\*\*75%)\*\*. \[\*\*GLM-4.7-GGUF\*\*\](https://huggingface.co/unsloth/GLM-4.7-GGUF) All uploads use Unsloth \[Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) for SOTA 5-shot MMLU and Aider performance, meaning you can run & fine-tune quantized GLM LLMs with minimal accuracy loss. ### :gear: Usage Guide The 2-bit dynamic quant UD-Q2\\\_K\\\_XL uses 135GB of disk space - this works well in a \*\*1x24GB card and 128GB of RAM\*\* with MoE offloading. The 1-bit UD-TQ1 GGUF also \*\*works natively in Ollama\*\*! {% hint style="info" %} You must use \`--jinja\` for llama.cpp quants - this uses our \[fixed chat templates\](#chat-template-bug-fixes) and enables the correct template! You might get incorrect results if you do not use \`--jinja\` {% endhint %} The 4-bit quants will fit in a 1x 40GB GPU (with MoE layers offloaded to RAM). Expect around 5 tokens/s with this setup if you have bonus 165GB RAM as well. It is recommended to have at least 205GB RAM to run this 4-bit. For optimal performance you will need at least 205GB unified memory or 205GB combined RAM+VRAM for 5+ tokens/s. To learn how to increase generation speed and fit longer contexts, \[read here\](#improving-generation-speed). {% hint style="success" %} Though not a must, for best performance, have your VRAM + RAM combined equal to the size of the quant you're downloading. If not, hard drive / SSD offloading will work with llama.cpp, just inference will be slower. Also use \`--fit on\` in \`llama.cpp\` to auto enable maximum GPU usage! {% endhint %} ### Recommended Settings Use distinct settings for different use cases. Recommended settings for default and multi-turn agentic use cases: | Default Settings (Most Tasks) | Terminal Bench, SWE Bench Verified | | ------------------------------------------------------------------ | ------------------------------------------------------------------ | | \*\*temperature = 1.0\*\* | \*\*temperature = 0.7\*\* | | \*\*top\\\_p = 0.95\*\* | \*\*top\\\_p = 1.0\*\* | | \`131072\` \*\*max new tokens\*\* | \`16384\` \*\*max new tokens\*\* | \* Use \`--jinja\` for llama.cpp variants - we \*\*fixed some chat template issues as well!\*\* \* \*\*Maximum context window:\*\* \`131,072\` ## Run GLM-4.7 Tutorials: See our step-by-step guides for running GLM-4.7 in \[Ollama\](#run-in-ollama) and \[llama.cpp\](#run-in-llama.cpp). ### ✨ Run in llama.cpp {% stepper %} {% step %} Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} If you want to use \`llama.cpp\` directly to load models, you can do the below: (:Q2\\\_K\\\_XL) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. Remember the model has only a maximum of 128K context length. \`\`\`bash export LLAMA\_CACHE="unsloth/GLM-4.7-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/GLM-4.7-GGUF:UD-Q2\_K\_XL \\ --jinja \\ --ctx-size 16384 \\ --flash-attn on \\ --temp 1.0 \\ --top-p 0.95 \\ --fit on \`\`\` {% hint style="info" %} \*\*Use \`--fit on\` introduced 15th Dec 2025 for maximum usage of your GPU and CPU.\*\* Optionally, try \`-ot ".ffn\_.\*\_exps.=CPU"\` to offload all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. And finally offload all layers via \`-ot ".ffn\_.\*\_exps.=CPU"\` This uses the least VRAM. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. {% endhint %} {% endstep %} {% step %} Download the model via (after installing \`pip install huggingface\_hub hf\_transfer\` ). You can choose \`UD-\`Q2\\\_K\\\_XL (dynamic 2bit quant) or other quantized versions like \`Q4\_K\_XL\` . We \*\*recommend using our 2.7bit dynamic quant\*\*\*\* \*\*\*\*\`UD-Q2\_K\_XL\`\*\*\*\* \*\*\*\*to balance size and accuracy\*\*. \`\`\`bash pip install -U huggingface\_hub hf download unsloth/GLM-4.7-GGUF \\ --local-dir unsloth/GLM-4.7-GGUF \\ --include "\*UD-Q2\_K\_XL\*" # Use "\*UD-TQ1\_0\*" for Dynamic 1bit \`\`\` {% endstep %} {% step %} You can edit \`--threads 32\` for the number of CPU threads, \`--ctx-size 16384\` for context length, \`--n-gpu-layers 2\` for GPU offloading on how many layers. Try adjusting it if your GPU goes out of memory. Also remove it if you have CPU only inference. {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/GLM-4.7-GGUF/UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ --jinja \\ --temp 1.0 \\ --top-p 0.95 \\ --ctx-size 16384 \\ --seed 3407 \\ --fit on \`\`\` {% endcode %} {% endstep %} {% endstepper %} ### :llama: Run in Ollama {% stepper %} {% step %} Install \`ollama\` if you haven't already! To run more variants of the model, \[see here\](https://unsloth.ai/docs/models/tutorials/pages/eK0BfjMHNrfvfe4HI6Pk#run-in-llama.cpp). \`\`\`bash apt-get update apt-get install pciutils -y curl -fsSL https://ollama.com/install.sh | sh \`\`\` {% endstep %} {% step %} Run the model! Note you can call \`ollama serve\`in another terminal if it fails! We include all our fixes and suggested parameters (temperature etc) in \`params\` in our Hugging Face upload! \`\`\`bash OLLAMA\_MODELS=unsloth ollama serve & OLLAMA\_MODELS=unsloth ollama run hf.co/unsloth/GLM-4.7-GGUF:TQ1\_0 \`\`\` {% endstep %} {% step %} To run other quants, you need to first merge the GGUF split files into 1 like the code below. Then you will need to run the model locally. \`\`\`bash ./llama.cpp/llama-gguf-split --merge \\ GLM-4.7-GGUF/GLM-4.7-UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ merged\_file.gguf \`\`\` \`\`\`bash OLLAMA\_MODELS=unsloth ollama serve & OLLAMA\_MODELS=unsloth ollama run merged\_file.gguf \`\`\` {% endstep %} {% endstepper %} ### ✨ Deploy with llama-server and OpenAI's completion library To use llama-server for deployment, use the following command: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\ --model unsloth/GLM-4.7-GGUF/UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ --alias "unsloth/GLM-4.7" \\ --fit on \\ --prio 3 \\ --temp 1.0 \\ --top-p 0.95 \\ --ctx-size 16384 \\ --port 8001 \\ --jinja \`\`\` {% endcode %} Then use OpenAI's Python library after \`pip install openai\` : \`\`\`python from openai import OpenAI import json openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/GLM-4.7", messages = \[{"role": "user", "content": "What is 2+2?"},\], ) print(completion.choices\[0\].message.content) \`\`\` ### :hammer:Tool Calling with GLM 4.7 See \[Tool Calling Guide\](/docs/basics/tool-calling-guide-for-local-llms.md) for more details on how to do tool calling. In a new terminal (if using tmux, use CTRL+B+D), we create some tools like adding 2 numbers, executing Python code, executing Linux functions and much more: {% code expandable="true" %} \`\`\`python import json, subprocess, random from typing import Any def add\_number(a: float | str, b: float | str) -> float: return float(a) + float(b) def multiply\_number(a: float | str, b: float | str) -> float: return float(a) \* float(b) def subtract\_number(a: float | str, b: float | str) -> float: return float(a) - float(b) def write\_a\_story() -> str: return random.choice(\[ "A long time ago in a galaxy far far away...", "There were 2 friends who loved sloths and code...", "The world was ending because every sloth evolved to have superhuman intelligence...", "Unbeknownst to one friend, the other accidentally coded a program to evolve sloths...", \]) def terminal(command: str) -> str: if "rm" in command or "sudo" in command or "dd" in command or "chmod" in command: msg = "Cannot execute 'rm, sudo, dd, chmod' commands since they are dangerous" print(msg); return msg print(f"Executing terminal command \`{command}\`") try: return str(subprocess.run(command, capture\_output = True, text = True, shell = True, check = True).stdout) except subprocess.CalledProcessError as e: return f"Command failed: {e.stderr}" def python(code: str) -> str: data = {} exec(code, data) del data\["\_\_builtins\_\_"\] return str(data) MAP\_FN = { "add\_number": add\_number, "multiply\_number": multiply\_number, "subtract\_number": subtract\_number, "write\_a\_story": write\_a\_story, "terminal": terminal, "python": python, } tools = \[ { "type": "function", "function": { "name": "add\_number", "description": "Add two numbers.", "parameters": { "type": "object", "properties": { "a": { "type": "string", "description": "The first number.", }, "b": { "type": "string", "description": "The second number.", }, }, "required": \["a", "b"\], }, }, }, { "type": "function", "function": { "name": "multiply\_number", "description": "Multiply two numbers.", "parameters": { "type": "object", "properties": { "a": { "type": "string", "description": "The first number.", }, "b": { "type": "string", "description": "The second number.", }, }, "required": \["a", "b"\], }, }, }, { "type": "function", "function": { "name": "subtract\_number", "description": "Subtract two numbers.", "parameters": { "type": "object", "properties": { "a": { "type": "string", "description": "The first number.", }, "b": { "type": "string", "description": "The second number.", }, }, "required": \["a", "b"\], }, }, }, { "type": "function", "function": { "name": "write\_a\_story", "description": "Writes a random story.", "parameters": { "type": "object", "properties": {}, "required": \[\], }, }, }, { "type": "function", "function": { "name": "terminal", "description": "Perform operations from the terminal.", "parameters": { "type": "object", "properties": { "command": { "type": "string", "description": "The command you wish to launch, e.g \`ls\`, \`rm\`, ...", }, }, "required": \["command"\], }, }, }, { "type": "function", "function": { "name": "python", "description": "Call a Python interpreter with some Python code that will be ran.", "parameters": { "type": "object", "properties": { "code": { "type": "string", "description": "The Python code to run", }, }, "required": \["code"\], }, }, }, \] \`\`\` {% endcode %} We then use the below functions (copy and paste and execute) which will parse the function calls automatically and call the OpenAI endpoint for any model: {% code overflow="wrap" expandable="true" %} \`\`\`python from openai import OpenAI def unsloth\_inference( messages, temperature = 0.7, top\_p = 0.95, top\_k = 40, min\_p = 0.01, repetition\_penalty = 1.0, ): messages = messages.copy() openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) model\_name = next(iter(openai\_client.models.list())).id print(f"Using model = {model\_name}") has\_tool\_calls = True original\_messages\_len = len(messages) while has\_tool\_calls: print(f"Current messages = {messages}") response = openai\_client.chat.completions.create( model = model\_name, messages = messages, temperature = temperature, top\_p = top\_p, tools = tools if tools else None, tool\_choice = "auto" if tools else None, extra\_body = {"top\_k": top\_k, "min\_p": min\_p, "repetition\_penalty" :repetition\_penalty,} ) tool\_calls = response.choices\[0\].message.tool\_calls or \[\] content = response.choices\[0\].message.content or "" tool\_calls\_dict = \[tc.to\_dict() for tc in tool\_calls\] if tool\_calls else tool\_calls messages.append({"role": "assistant", "tool\_calls": tool\_calls\_dict, "content": content,}) for tool\_call in tool\_calls: fx, args, \_id = tool\_call.function.name, tool\_call.function.arguments, tool\_call.id out = MAP\_FN\[fx\](\*\*json.loads(args)) messages.append({"role": "tool", "tool\_call\_id": \_id, "name": fx, "content": str(out),}) else: has\_tool\_calls = False return messages \`\`\` {% endcode %} After launching GLM 4.7 via \`llama-server\` like in \[#deploy-with-llama-server-and-openais-completion-library\](#deploy-with-llama-server-and-openais-completion-library "mention") or see \[Tool Calling Guide\](/docs/basics/tool-calling-guide-for-local-llms.md) for more details, we then can do some tool calls: \*\*Tool Call for mathematical operations for GLM 4.7\*\* {% code overflow="wrap" %} \`\`\`python messages = \[{ "role": "user", "content": \[{"type": "text", "text": "What is today's date plus 3 days?"}\], }\] unsloth\_inference(messages, temperature = 0.7, top\_p = 1.0, top\_k = -1, min\_p = 0.00) \`\`\` {% endcode %} ![](https://unsloth.ai/files/kGn0zF1WiJINokTkwJXz) \*\*Tool Call to execute generated Python code for GLM 4.7\*\* {% code overflow="wrap" %} \`\`\`python messages = \[{ "role": "user", "content": \[{"type": "text", "text": "Create a Fibonacci function in Python and find fib(20)."}\], }\] unsloth\_inference(messages, temperature = 0.7, top\_p = 1.0, top\_k = -1, min\_p = 0.00) \`\`\` {% endcode %} ![](https://unsloth.ai/files/N11O7CXrykHBLmJaCn6Q) \### :snowboarder: Improving generation speed \*\*Use \`--fit on\` introduced 15th Dec 2025 for maximum usage of your GPU and CPU. See\*\* \[\*\*https://github.com/ggml-org/llama.cpp/pull/16653\*\*\](https://github.com/ggml-org/llama.cpp/pull/16653) \*\*\`--fit on\` auto offloads as much of the model as possible to the GPU, then places the rest on CPU.\*\* If you have more VRAM, you can try offloading more MoE layers, or offloading whole layers themselves. Normally, \`-ot ".ffn\_.\*\_exps.=CPU"\` offloads all MoE layers to the CPU! This effectively allows you to fit all non MoE layers on 1 GPU, improving generation speeds. You can customize the regex expression to fit more layers if you have more GPU capacity. If you have a bit more GPU memory, try \`-ot ".ffn\_(up|down)\_exps.=CPU"\` This offloads up and down projection MoE layers. Try \`-ot ".ffn\_(up)\_exps.=CPU"\` if you have even more GPU memory. This offloads only up projection MoE layers. You can also customize the regex, for example \`-ot "\\.(6|7|8|9|\[0-9\]\[0-9\]|\[0-9\]\[0-9\]\[0-9\])\\.ffn\_(gate|up|down)\_exps.=CPU"\` means to offload gate, up and down MoE layers but only from the 6th layer onwards. Llama.cpp also introduces high throughput mode. Use \`llama-parallel\`. Read more about it \[here\](https://github.com/ggml-org/llama.cpp/tree/master/examples/parallel). You can also \*\*quantize the KV cache to 4bits\*\* for example to reduce VRAM / RAM movement, which can also make the generation process faster. ### 📐How to fit long context (full 128K) To fit longer context, you can use \*\*KV cache quantization\*\* to quantize the K and V caches to lower bits. This can also increase generation speed due to reduced RAM / VRAM data movement. The allowed options for K quantization (default is \`f16\`) include the below. \`--cache-type-k f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` You should use the \`\_1\` variants for somewhat increased accuracy, albeit it's slightly slower. For eg \`q4\_1, q5\_1\` You can also quantize the V cache, but you will need to \*\*compile llama.cpp with Flash Attention\*\* support via \`-DGGML\_CUDA\_FA\_ALL\_QUANTS=ON\`, and use \`--flash-attn\` to enable it. Then you can use together with \`--cache-type-k\` : \`--cache-type-v f32, f16, bf16, q8\_0, q4\_0, q4\_1, iq4\_nl, q5\_0, q5\_1\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/tutorials/glm-4.7.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/nouveau/studio/chat.md). # Comment exécuter des modèles avec Unsloth Studio \[Unsloth Studio\](/docs/fr/nouveau/studio.md) vous permet d'exécuter des modèles d'IA 100 % hors ligne sur votre ordinateur. Exécutez des formats de modèles comme GGUF et safetensors depuis Hugging Face ou depuis vos fichiers locaux. \* \*\*Fonctionne sur toutes les configurations MacOS, CPU, Windows, Linux et WSL ! Aucun GPU requis\*\* \* \[\*\*Appel d’outils auto-réparateur\*\*\](#auto-healing-tool-calling)\*\*,\*\* avancé \[\*\*recherche web\*\*\](#advanced-web-search), \[\*\*de code\*\*\](#code-execution) \* Utilisez Unsloth comme une inférence compatible avec OpenAI \[\*\*point de terminaison d'API\*\*\](/docs/fr/notions-de-base/api.md) ou connectez un \[fournisseur\](/docs/fr/integrations/connections.md) \* Rechercher + Télécharger + Exécuter + \[Comparer\](#model-arena) n'importe quel modèle comme des GGUF, des adaptateurs LoRA, des safetensors, etc. \* \[\*\*Réglage automatique des paramètres d'inférence\*\*\](#auto-parameter-tuning) (temp, top-p, etc.) et modification des gabarits de chat \* Téléversez des images, de l'audio, des PDF, du code, des DOCX et d'autres types de fichiers pour discuter avec. ![](https://unsloth.ai/files/5b73218a32955029943d07713e050847da7e3e01) \### Utilisation de Unsloth Studio Chat {% hint style="success" %} Unsloth Studio Chat fonctionne automatiquement sur \*\*les configurations multi-GPU\*\* pour l'inférence. {% endhint %} {% columns %} {% column %} #### Exécution de code Unsloth Studio permet aux LLM d'exécuter Bash et Python, pas seulement JavaScript. Il met aussi en sandbox des programmes comme Claude Artifacts afin que les modèles puissent tester du code, générer des fichiers et vérifier les réponses avec de vrais calculs. Cela rend les réponses des modèles plus fiables et plus précises. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/8d3032a7bd41a3a58d8581c0b9c8febe68e48077) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Appel d'outils auto-réparateur Unsloth Studio permet non seulement \[appels d'outils\](#id-50-tool-calling-accuracy), mais corrige aussi automatiquement les appels d'outils mal formés ou cassés de 50 %. Cela signifie que vous obtiendrez toujours des sorties d'inférence \*\*sans\*\* appels d'outils défectueux. Par exemple, Qwen3.5-4B a recherché plus de 20 sites web et cité ses sources, la recherche web se déroulant à l'intérieur de sa trace de réflexion. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/b03f506381fed38c517bc09e2e726bc0013f23f4) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Recherche web avancée La recherche web d'Unsloth visite réellement les pages directement pour collecter des informations et des données pertinentes, et ne se contente pas de parcourir les résumés des sites web. Cela fournit des sorties beaucoup plus précises / des informations et un contexte plus approfondis. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/5b73218a32955029943d07713e050847da7e3e01) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Utilisez Unsloth comme point de terminaison d'API Vous pouvez désormais utiliser des LLM locaux via des outils comme \[Claude Code\](/docs/fr/notions-de-base/claude-code.md) et \[Codex\](/docs/fr/notions-de-base/codex.md) en le connectant au \[point de terminaison d'API\](#use-unsloth-as-an-api-endpoint). Cela signifie que vous pourrez exécuter directement des modèles Qwen et Gemma dans ces outils avec l'inférence d'Unsloth, qui inclut des fonctionnalités comme les appels d'outils auto-réparateurs, la recherche web, etc. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/1a2d152a014c5c542c774dac8c97d657a9f4124f) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Paramètres d'inférence automatiques Les paramètres d'inférence comme \*\*la température\*\*, \*\*top-p\*\*, \*\*top-k\*\*, \[\*\*MTP\*\*\](/docs/fr/modeles/qwen3.6.md#mtp-guide) sont automatiquement prédéfinis pour les nouveaux modèles comme Qwen3.5 afin que vous obteniez les meilleurs résultats sans vous soucier des réglages. Vous pouvez aussi ajuster les paramètres manuellement et modifier le prompt système. L'ajustement de la longueur du contexte n'est plus nécessaire avec le contexte automatique intelligent de llama.cpp, qui n'utilise que le contexte dont vous avez besoin sans charger quoi que ce soit en plus. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/85fc8084702579bd11c56183ae0b99d11cafcf65) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Connecter des fournisseurs \[Unsloth se connecte\](/docs/fr/integrations/connections.md) à OpenAI, Anthropic, Ollama, llama.cpp, vLLM et autres. Ajoutez des clés API ou des URL de serveur de modèles, puis utilisez des modèles externes dans la même interface de chat que les modèles locaux + cloud. Exécutez avec \[la mise en cache des prompts\](/docs/fr/integrations/connections.md#prompt-caching), les appels d'outils, le raisonnement et des fonctionnalités natives du fournisseur comme celles d'OpenAI \[recherche web\](#web-search-and-thinking) et \[de code\](#code-execution). {% endcolumn %} {% column %} ![](https://unsloth.ai/files/85fc8084702579bd11c56183ae0b99d11cafcf65) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Rechercher et exécuter des modèles Vous pouvez rechercher et télécharger n'importe quel modèle via Hugging Face ou utiliser des fichiers locaux. Unsloth prend en charge un large éventail de types de modèles, notamment \*\*GGUF\*\*, de vision-langage et de synthèse vocale. Exécutez les derniers modèles comme \[Qwen3.5\](/docs/fr/modeles/qwen3.5.md) ou NVIDIA \[Nemotron 3\](/docs/fr/modeles/nemotron-3.md). Téléversez des images, de l'audio, des PDF, du code, des DOCX et d'autres types de fichiers pour discuter avec. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/2885112f766f1614a56fe25756dd558b131aa3f2) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Espace de travail de chat Saisissez des invites, joignez n'importe quels documents, images (webp, png), fichiers de code, txt ou audio comme contexte supplémentaire, et voyez les réponses du modèle en temps réel. Activez ou désactivez : Raisonnement + recherche web. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/2f8142ea99489385496d6332e8bdf629d599584f) {% endcolumn %} {% endcolumns %} ### \*\*+50 % de précision des appels d'outils\*\* Unsloth offre plusieurs fonctionnalités uniques qui améliorent les appels d'outils, notamment : \* Les appels d'outils sur tous les modèles dans Unsloth sont \*\*30 % à 80 % plus précis\*\*. \* La recherche web récupère le contenu web réel au lieu de simples résumés. \* Le nombre maximal d'appels d'outils autorisés est \*\*supérieur à 25.\*\* \* Les appels d'outils se terminent de manière plus fiable, ce qui réduit les boucles et les appels répétés. \* Une logique améliorée de réparation des appels d'outils et de déduplication aide à empêcher les fuites de XML dans les sorties. Voir les résultats des tests avec \`unsloth/Qwen3.5-4B-GGUF (UD-Q4\_K\_XL)\` avec la recherche web, l'exécution de code et le raisonnement activés : | Métrique | Appel d'outils normal | Appel d'outils Unsloth | | ----------------------------------------- | --------------------- | ---------------------- | | Fuites XML dans la réponse | 10/10 | 0/10 | | Récupérations d'URL utilisées | 0 | 4/10 exécutions | | Exécutions avec les bons noms de chansons | 0/10 | 2/10 | | Nombre moyen d'appels d'outils | 5.5 | 3.8 | | Temps de réponse moyen | 12,3 s | 9,8 s | ### Arène des modèles Unsloth Chat vous permet de comparer côte à côte n'importe quels deux modèles en utilisant la même invite. Par exemple, comparez le modèle de base et l'adaptateur LoRA. L'inférence chargera d'abord un modèle, puis le second (l'inférence parallèle est en cours de développement). ![](https://unsloth.ai/files/a2e0d4bfd76d0287d9b02f802fd667c9e1ac821e) {% columns %} {% column %} Après l'entraînement, vous pouvez comparer côte à côte le modèle de base et le modèle affiné avec la même invite pour voir ce qui a changé et si les résultats se sont améliorés. Ce flux de travail facilite la visualisation de la manière dont votre fine-tuning a modifié les réponses du modèle et si cela a amélioré les résultats pour votre cas d'usage. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/363f22dc6187049e754bdae94b495a45c68bb1b9) {% endcolumn %} {% endcolumns %} {% hint style="success" %} Unsloth Studio Chat fonctionne automatiquement sur \*\*les configurations multi-GPU\*\* pour l'inférence. {% endhint %} ### Utilisation de modèles GGUF anciens / existants {% columns %} {% column %} \*\*Mise à jour du 1er avril :\*\* Vous pouvez maintenant sélectionner un dossier existant pour qu'Unsloth le détecte. \*\*Mise à jour du 27 mars :\*\* Unsloth Studio détecte désormais \*\*automatiquement les modèles anciens / préexistants\*\* téléchargés depuis Hugging Face, LM Studio, etc. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/089d4b6412408ae0b8fd66eecc20eac7aa886d13) {% endcolumn %} {% endcolumns %} \*\*Instructions manuelles :\*\* Unsloth Studio détecte les modèles téléchargés dans votre cache Hugging Face Hub \`(C:\\Users{your\_username}.cache\\huggingface\\hub)\`. Si vous avez téléchargé des modèles GGUF via LM Studio, notez qu'ils sont stockés dans \`C:\\Users\\{your\_username}.cache\\lm-studio\\models\` \*\*\*OU\*\*\* \`C:\\Users{your\_username}\\lm-studio\\models\` et ne sont pas visibles par défaut pour llama.cpp - vous devrez déplacer ou copier ces fichiers .gguf dans votre répertoire de cache Hugging Face Hub (ou dans un autre chemin accessible à llama.cpp) pour qu'Unsloth Studio puisse les charger. Après avoir affiné un modèle ou un adaptateur dans Unsloth, vous pouvez l'exporter en GGUF et exécuter une inférence locale avec \*\*llama.cpp\*\* directement dans Unsloth Chat. Unsloth Studio est propulsé par llama.cpp et Hugging Face. ### Ajout de fichiers comme contexte Unsloth Chat prend en charge les entrées multimodales directement dans la conversation. Vous pouvez joindre des documents, des images ou de l'audio comme contexte supplémentaire pour une invite. ![](https://unsloth.ai/files/23726c6bd636565d1a3e66276e12b90752273f88) Cela facilite les tests de la manière dont un modèle gère des entrées réelles telles que des PDF, des captures d'écran ou du matériel de référence. Les fichiers sont traités localement et inclus comme contexte pour le modèle. ### \*\*Suppression des fichiers de modèle\*\* Vous pouvez supprimer d'anciens fichiers de modèle soit depuis l'icône de corbeille dans la recherche de modèles, soit en supprimant le dossier de modèle mis en cache correspondant dans le répertoire de cache Hugging Face par défaut. Par défaut, Hugging Face utilise \`~/.cache/huggingface/hub/\` sur macOS/Linux/WSL et \`C:\\Users\\\\.cache\\huggingface\\hub\\\` sur Windows. \* \*\*MacOS, Linux, WSL :\*\* \`~/.cache/huggingface/hub/\` \* \*\*Windows :\*\* \`%USERPROFILE%\\.cache\\huggingface\\hub\\\` Si \`HF\_HUB\_CACHE\` ou \`HF\_HOME\` est défini, utilisez cet emplacement à la place. Sur Linux et WSL, \`XDG\_CACHE\_HOME\` peut également modifier la racine de cache par défaut. ### \*\*Unsloth ne détecte pas ou n'utilise pas mon GPU\*\* Si le modèle n'utilise pas votre GPU, en particulier avec Docker, essayez : Téléchargement manuel de la dernière image : \`\`\`bash docker pull unsloth/unsloth:latest \`\`\` \* Démarrez le conteneur avec accès GPU : \* \`docker run\`: \`--gpus all\` \* Docker Compose : \`capabilities: \[gpu\]\` \* Sur Linux, assurez-vous que NVIDIA Container Toolkit est installé. \* Sur Windows : \* Vérifiez que \`nvcc --version\` correspond à la version de CUDA affichée dans \`nvidia-smi\` \* Suivez : \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/nouveau/studio/chat.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md). # Reinforcement Learning visuel (VLM RL) Unsloth prend désormais en charge le RL vision/multimodal avec \[Qwen3-VL\](/docs/fr/modeles/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md), \[Gemma 3\](/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md) et plus encore. En raison du \[partage de poids\](/docs/fr/commencer/reinforcement-learning-rl-guide.md#what-unsloth-offers-for-rl) et des noyaux personnalisés d'Unsloth, Unsloth rend le RL VLM \*\*1,5–2× plus rapide,\*\* utilise \*\*90% moins de VRAM\*\*, et permet \*\*des contextes\*\* 15× plus longs que les configurations FA2, sans perte de précision. Cette mise à jour introduit également l' \[GSPO\](#gspo-rl) algorithme. Unsloth peut entraîner Qwen3-VL-8B avec GSPO/GRPO sur un GPU Colab T4 gratuit. D'autres VLM fonctionnent aussi, mais peuvent nécessiter des GPU plus grands. Gemma exige des GPU plus récents que le T4 parce que vLLM \[se limite à Bfloat16\](/docs/fr/modeles/tutorials/gemma-3-how-to-run-and-fine-tune.md#unsloth-fine-tuning-fixes), nous recommandons donc NVIDIA L4 sur Colab. Nos notebooks résolvent des problèmes de mathématiques numériques impliquant des images et des schémas : \* \*\*Qwen-3 VL-8B\*\* (inférence vLLM)\*\*:\*\* \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) \* \*\*Qwen-2.5 VL-7B\*\* (inférence vLLM)\*\*:\*\* \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2\_5\_7B\_VL\_GRPO.ipynb) •\[ Kaggle\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen2\_5\_7B\_VL\_GRPO.ipynb\\&accelerator=nvidiaTeslaT4) \* \*\*Gemma-3-4B\*\* (inférence Unsloth) : \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision-GRPO.ipynb) Nous avons également intégré nativement vLLM VLM dans Unsloth, donc tout ce que vous avez à faire pour utiliser l'inférence vLLM est d'activer le \`fast\_inference=True\` lors de l'initialisation du modèle. Remerciements particuliers à \[Sinoué GAD\](https://github.com/unslothai/unsloth/pull/2752) pour avoir fourni le \[premier notebook\](https://github.com/GAD-cell/vlm-grpo/blob/main/examples/VLM\_GRPO\_basic\_example.ipynb) qui a rendu l'intégration du RL VLM plus facile ! Ce support VLM intègre également notre dernière mise à jour pour un RL encore plus économe en mémoire et plus rapide, incluant notre \[fonctionnalité Standby\](/docs/fr/commencer/reinforcement-learning-rl-guide/memory-efficient-rl.md#unsloth-standby), qui limite de manière unique la dégradation de la vitesse par rapport à d'autres implémentations. {% hint style="info" %} Vous ne pouvez utiliser que \`fast\_inference\` pour les VLM pris en charge par vLLM. Certains modèles, comme Llama 3.2 Vision, ne peuvent donc fonctionner qu'en dehors de vLLM, mais ils fonctionnent toujours dans Unsloth. {% endhint %} \`\`\`python os.environ\['UNSLOTH\_VLLM\_STANDBY'\] = '1' # Pour activer GRPO économe en mémoire avec vLLM model, tokenizer = FastVisionModel.from\_pretrained( model\_name = "Qwen/Qwen2.5-VL-7B-Instruct", max\_seq\_length = 16384, # Doit être aussi grand pour insérer l'image dans le contexte load\_in\_4bit = True, # False pour LoRA 16bit fast\_inference = True, # Activer l'inférence rapide vLLM gpu\_memory\_utilization = 0.8, # Réduire si mémoire insuffisante ) \`\`\` Il est également important de noter que vLLM ne prend pas en charge LoRA pour les couches vision/encodeur, définissez donc \`finetune\_vision\_layers = False\` lors du chargement d'un adaptateur LoRA.\\ Cependant, vous POUVEZ entraîner également les couches vision si vous utilisez l'inférence via transformers/Unsloth. \`\`\`python # Ajouter l'adaptateur LoRA au modèle pour un ajustement fin efficace en paramètres model = FastVisionModel.get\_peft\_model( model, finetune\_vision\_layers = False,# fast\_inference ne prend pas encore en charge finetune\_vision\_layers :( finetune\_language\_layers = True, # False si vous n'affinez pas les couches de langage finetune\_attention\_modules = True, # False si vous n'affinez pas les couches d'attention finetune\_mlp\_modules = True, # False si vous n'affinez pas les couches MLP r = lora\_rank, # Choisissez n'importe quel nombre > 0 ! Suggestions : 8, 16, 32, 64, 128 lora\_alpha = lora\_rank\*2, # \*2 accélère l'entraînement use\_gradient\_checkpointing = "unsloth", # Réduit l'utilisation de la mémoire random\_state = 3407, ) \`\`\` ## :butterfly:Problèmes et particularités du RL Vision Qwen 2.5 VL Pendant le RL pour Qwen 2.5 VL, vous pourriez voir la sortie d'inférence suivante : {% code overflow="wrap" %} \`\`\` addCriterion \\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n\\n addCriterion\\n\\n 自动生成\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n\\n addCriterion\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n \`\`\` {% endcode %} Ceci a été \[signalé\](https://github.com/QwenLM/Qwen2.5-VL/issues/759) également dans Qwen2.5-VL-7B-Instruct comme résultat inattendu « addCriterion ». En fait nous voyons cela aussi ! Nous avons essayé des machines non Unsloth, bfloat16 et float16 et d'autres choses, mais cela semble persister. Par exemple l'élément 165 c.-à-d. \`train\_dataset\[165\]\` du \[AI4Math/MathVista\](https://huggingface.co/datasets/AI4Math/MathVista) jeu de données est ci-dessous : {% code overflow="wrap" %} \`\`\` La figure est une vue aérienne du trajet emprunté par un pilote de course lorsque sa voiture heurte le mur de la piste. Juste avant la collision, il se déplace à la vitesse v\_i = 70 \\mathrm{~m} / \\mathrm{s} le long d'une ligne droite à 30^{\\circ} du mur. Juste après la collision, il se déplace à la vitesse v\_f = 50 \\mathrm{~m} / \\mathrm{s} le long d'une ligne droite à 10^{\\circ} du mur. Sa masse m est de 80 \\mathrm{~kg}. La collision dure 14 \\mathrm{~ms}. Quelle est la magnitude de la force moyenne sur le pilote pendant la collision ? \`\`\` {% endcode %} ![](https://unsloth.ai/files/aef3082644571b3c158bc53aeac3a1c65c100bc7) Et ensuite nous obtenons la sortie incompréhensible ci‑dessus. On pourrait ajouter une fonction de récompense pour pénaliser l'ajout de addCriterion, ou pénaliser les sorties incompréhensibles. Cependant, l'autre approche est de l'entraîner plus longtemps. Par exemple, seulement après environ 60 étapes nous voyons le modèle réellement apprendre via le RL : ![](https://unsloth.ai/files/634c0d6aab4a66b6c138e3aaa625ce81254ea7a7) {% hint style="success" %} Forcer \`<|assistant|>\` pendant la génération réduira les occurrences de ces résultats incompréhensibles comme prévu puisque c'est un modèle Instruct, cependant il est toujours préférable d'ajouter une fonction de récompense pour pénaliser les mauvaises générations, comme décrit dans la section suivante. {% endhint %} ## :medal:Fonctions de récompense pour réduire les sorties incompréhensibles Pour pénaliser \`addCriterion\` et les sorties incompréhensibles, nous avons modifié la fonction de récompense pour pénaliser trop de \`addCriterion\` et de sauts de ligne. \`\`\`python def formatting\_reward\_func(completions,\*\*kwargs): import re thinking\_pattern = f'{REASONING\_START}(.\*?){REASONING\_END}' answer\_pattern = f'{SOLUTION\_START}(.\*?){SOLUTION\_END}' scores = \[\] for completion in completions: score = 0 thinking\_matches = re.findall(thinking\_pattern, completion, re.DOTALL) answer\_matches = re.findall(answer\_pattern, completion, re.DOTALL) if len(thinking\_matches) == 1: score += 1.0 if len(answer\_matches) == 1: score += 1.0 # Corriger les problèmes addCriterion # Voir https://docs.unsloth.ai/new/vision-reinforcement-learning-vlm-rl#qwen-2.5-vl-vision-rl-issues-and-quirks # Pénaliser l'excès de addCriterion et de sauts de ligne if len(completion) != 0: removal = completion.replace("addCriterion", "").replace("\\n", "") if (len(completion)-len(removal))/len(completion) >= 0.5: score -= 2.0 scores.append(score) return scores \`\`\` ## :checkered\\\_flag:Reinforcement Learning GSPO Cette mise à jour ajoute en outre GSPO (\[Group Sequence Policy Optimization\](https://arxiv.org/abs/2507.18071)) qui est une variante de GRPO créée par l'équipe Qwen d'Alibaba. Ils ont remarqué que GRPO implique implicitement des poids d'importance pour chaque token, même si les avantages explicites ne se mettent pas à l'échelle ou ne changent pas avec chaque token. Cela a conduit à la création de GSPO, qui assigne maintenant l'importance sur la vraisemblance de la séquence plutôt que sur les vraisemblances individuelles des tokens. La différence entre ces deux algorithmes peut être vue ci‑dessous, issue à la fois de l'article GSPO de Qwen et Alibaba : ![](https://unsloth.ai/files/71070c1f23656c6ea3da36c262f7e943361cc5a0) Algorithme GRPO, Source : [Qwen](https://arxiv.org/abs/2507.18071) ![](https://unsloth.ai/files/2c91908097ff4f0243bbcc6092be9984f0807a9a) Algorithme GSPO, Source : [Qwen](https://arxiv.org/abs/2507.18071) Dans l'équation 1, on peut voir que les avantages mettent à l'échelle chacune des lignes dans les logprobs des tokens avant que ce tenseur ne soit sommée. Essentiellement, chaque token reçoit la même mise à l'échelle bien que cette mise à l'échelle ait été appliquée à l'ensemble de la séquence plutôt qu'à chaque token individuel. Un diagramme simple de ceci peut être vu ci‑dessous : ![](https://unsloth.ai/files/a07e036e2ce9fa0e3917825b78308e78e3233bec) Ratio de logprob GRPO mis à l'échelle ligne par ligne avec les avantages L'équation 2 montre que les ratios de logprob pour chaque séquence sont sommés et exponentiés après le calcul des ratios de logprob, et seuls les ratios de séquence résultants sont multipliés ligne par ligne par les avantages. ![](https://unsloth.ai/files/1ec6e15370ae70f72c4d6ff4359cbe2f2db08d84) Ratio de séquence GSPO mis à l'échelle ligne par ligne avec les avantages Activer GSPO est simple, il vous suffit de définir le \`importance\_sampling\_level = "sequence"\` indicateur dans la configuration GRPO. \`\`\`python training\_args = GRPOConfig( output\_dir = "vlm-grpo-unsloth", per\_device\_train\_batch\_size = 8, gradient\_accumulation\_steps = 4, learning\_rate = 5e-6, adam\_beta1 = 0.9, adam\_beta2 = 0.99, weight\_decay = 0.1, warmup\_ratio = 0.1, lr\_scheduler\_type = "cosine", optim = "adamw\_8bit", # beta = 0.00, epsilon = 3e-4, epsilon\_high = 4e-4, num\_generations = 8, max\_prompt\_length = 1024, max\_completion\_length = 1024, log\_completions = False, max\_grad\_norm = 0.1, temperature = 0.9, # report\_to = "none", # Mettre à "wandb" si vous souhaitez enregistrer sur Weights & Biases num\_train\_epochs = 2, # Pour un test rapide, augmenter pour un entraînement complet report\_to = "none" # GSPO est ci‑dessous : importance\_sampling\_level = "sequence", # Dr GRPO / GAPO etc loss\_type = "dr\_grpo", ) \`\`\` Dans l'ensemble, Unsloth permet désormais avec l'inférence rapide VLM vLLM à la fois une réduction de 90% de l'utilisation de la mémoire mais aussi une vitesse 1,5–2x plus élevée avec GRPO et GSPO ! Si vous souhaitez en savoir plus sur le reinforcement learning, consultez notre guide RL : \[Reinforcement Learning\](/docs/fr/commencer/reinforcement-learning-rl-guide.md) \*\*\*Auteurs :\*\* Un grand merci à\* \[\*Keith\*\](https://www.linkedin.com/in/keith-truongcao-7bb84a23b/) \*et\* \[\*Datta\*\](https://www.linkedin.com/in/datta0/) \*pour avoir contribué à cet article !\* --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context.md). # Reinforcement Learning GRPO with 7x Longer Context Reinforcement learning's (RL) biggest challenge is supporting long reasoning traces. We're introducing new batching algorithms to enable \\~\*\*7x longer context\*\* (can be more than 12x) RL training with no accuracy or speed degradation vs. other optimized setups that use FA3, kernels & chunked losses. \* Unsloth now trains gpt-oss QLoRA with \*\*380K context\*\* on a single 192GB NVIDIA B200 GPU \* \[Qwen3\](/docs/models/tutorials/qwen3-how-to-run-and-fine-tune.md#fine-tuning-qwen3-with-unsloth)-8B GRPO reaches \*\*110K context\*\* on an 80GB VRAM H100 via \[vLLM\](#vllm-for-rl) and QLoRA, and \*\*65K\*\* for \[gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune/gpt-oss-reinforcement-learning.md) with BF16 LoRA. \* On 24GB VRAM, gpt-oss reaches 20K context and 32K for \[Qwen3-VL\](/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md)-8B QLoRA \* Unsloth GRPO RL runs with Llama, Gemma & all models auto support longer contexts Our new data-movement and batching kernels and algorithms unlocks more context by: \* Dynamic \[flattened sequence chunking\](#flattened-sequence-length-chunking) to avoid materializing massive logit tensors and \* \[Offloading log softmax\](#offloading-activations-for-log-softmax) activations which prevents silent memory growth over time. {% hint style="info" %} \*\*You can combine all features in Unsloth together:\*\* 1. Unsloth's \[weight-sharing\](/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl.md) feature with \[vLLM\](https://github.com/vllm-project/vllm) and our Standby Feature in \[Memory Efficient RL\](/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl.md) 2. Unsloth's \[Flex Attention\](/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md) for long context gpt-oss and our \[500K Context Training\](/docs/blog/500k-context-length-fine-tuning.md) 3. Float8 training in \[FP8 RL\](/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning.md) and Unsloth's \[async gradient checkpointing\](https://unsloth.ai/blog/long-context) and much more {% endhint %} ### :tada:Getting started To get started, you can use any existing \[GRPO notebooks\](/docs/get-started/unsloth-notebooks.md#grpo-reasoning-rl-notebooks) (or update Unsloth if local): {% columns %} {% column width="33.33333333333333%" %} \[\*\*gpt-oss-20b\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) GSPO {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[\*\*Qwen3-VL-8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) Vision RL {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[Qwen3-8B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_8B\_FP8\_GRPO.ipynb) L4 GPU {% embed url="" %} {% endcolumn %} {% endcolumns %} Adopting Unsloth for your RL tasks provides a robust framework for managing large-scale models efficiently. To effectively utilize Unsloth's enhancements: \* \*\*Hardware Recommendations\*\*: Use of NVIDIA H100 or equivalent for optimal VRAM utilization. \* \*\*Configuration Tips\*\*: Ensure \`batch\_size\` and \`gradient\_accumulation\_steps\` settings align with your computational resources for best performance. {% hint style="success" %} Update Unsloth to the latest Pypi release to get the latest updates: \`\`\` pip install --upgrade --no-cache-dir unsloth unsloth\_zoo \`\`\` {% endhint %} Our benchmarks highlight the memory savings achieved in comparison to earlier versions for GPT OSS and Qwen3-8B. Both plots below (without \[standby\](/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl.md)) were run with \`batch\_size = 4\` and \`gradient\_accumulation\_steps=2\` , since standby by design uses all VRAM. For our benchmarks, we compare BF16 GRPO to Hugging Face with all optimizations enabled (all kernels in kernels library, Flash Attention 3, chunked loss kernels, etc): ### :1234:Flattened sequence length chunking Previously, Unsloth reduced memory usage of RL by avoiding the full materialization of the logits tensor through chunking over the batch dimension. A rough estimate of the VRAM required to materialize logits during the forward pass is shown in Equation (1). $$ \\text{Equation 1: } \\text{Logit Memory (GB)} = \\frac{\\text{batch size} \\times\\text{context length} \\times \\text{vocab dim}}{1024^3} $$ Using this formulation, a configuration with \`batch\_size = 4\`, \`context\_length = 8192\`, and \`vocab\_dim = 128,000\` would require approximately \*\*3.3 GB of VRAM\*\* to store the logits tensor. Via \[Long Context gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md) last year, we then introduced a fused loss approach for GRPO. This approach ensures that only a single batch sample is processed at a time, significantly reducing peak memory usage. Under the same configuration, VRAM usage drops to approximately \*\*0.83 GB\*\*, as reflected in Equation (2). $$ \\text{Equation 2: }\\text{Logit Memory (GB)} = \\frac{\\text{context length} \\times \\text{vocab dim}}{1024^3} $$ ![](https://unsloth.ai/files/0YRu9DBGYHq4OgfK9jzs) Figure 1: gpt-oss BF16 GRPO LoRA (Unsloth vs. HF with all optimizations on) ![](https://unsloth.ai/files/u1OAtTBbsIwyaR4QJlU2) Figure 2: Qwen3-8B QLoRA GRPO LoRA (Unsloth vs. HF with all optimizations on) In this update, we extend the same idea further by introducing chunking across the \*\*sequence dimension\*\* as well. Instead of materializing logits for the entire \`(batch\_size × context\_length)\` space at once, we flatten these dimensions and process them in smaller chunks using a configurable multiplier. This allows Unsloth to support substantially longer contexts without increasing peak memory usage. In Figure 5 below, we use a multiplier of \`max(4, context\_length // 4096)\`, though any multiplier can be specified depending on the desired memory–performance tradeoff. With this setting, the same example configuration (\`batch\_size = 4\`, \`context\_length = 8192\`, \`vocab\_dim = 128,000\`) now requires only \*\*0.207 GB of VRAM\*\* for logits materialization. $$ \\text{Equation 3: }\\text{Logit Memory (GB)} = \\frac{\\frac{\\text{context length}}{\\text{multiplier}} \\times \\text{vocab dim}}{1024^3} $$ ![](https://unsloth.ai/files/U2qj0dGVLOAkAKrrgDPA) Figure 3: gpt-oss-20b (H100) Unsloth new vs. old ![](https://unsloth.ai/files/qAp6wYuqD7eljedEQlV5) Figure 4: Qwen3-8B (H100) Unsloth new vs. old ![](https://unsloth.ai/files/IfYyM8hgKyQHV0szmgti) Figure 5: gpt-oss-20b (H100) ![](https://unsloth.ai/files/PrZFqzlOZjOHGilFNaBn) Figure 6: Qwen3-8B (B200) This update is reflected in the compiled \`chunked\_hidden\_states\_selective\_log\_softmax\` below, which now supports chunking across both the batch and sequence dimensions. To preserve the logits tensor (\`\[batch\_size, context\_length, vocab\_dim\]\`), it is always chunked across the batch dimension. Additional sequence chunking is controlled via \`unsloth\_logit\_chunk\_multiplier\` in the GRPO configuration; if unset, it defaults to \`max(4, context\_length // 4096)\`. In the example below, \`input\_ids\_chunk\[0\]\` corresponds to the size of the hidden states mini batches in optimization 2. \`\`\`python logprobs\_chunk = chunked\_hidden\_states\_selective\_log\_softmax( new\_hidden\_states\_chunk, lm\_head, completion\_ids, chunks=input\_ids\_chunk.shape\[0\]\*multiplier, logit\_scale\_multiply=logit\_scale\_multiply, logit\_scale\_divide=logit\_scale\_divide, logit\_softcapping=logit\_softcapping, temperature=temperature, ) \`\`\` 1. We utilize torch.compile with custom compile options to reduce VRAM and increase speed. 2. All chunked logits are upcasted in float32 to preserve accuracy. 3. We support logit softcapping, temperature scaling and all other features. ### :ghost:Hidden States Chunking We also observed that at longer context lengths, hidden states can become a significant contributor to memory usage. For demonstration, we will assume \`hidden\_states\_dim=4096\`. The corresponding memory usage follows a similar formulation to the logits case, shown below. $$ \\text{Hidden States Memory (GB)} = \\frac{\\text{batch size} \\times\\text{context length} \\times \\text{hidden states dim}}{1024^3} $$ With a \`batch\_size = 8\` and \`context\_length = 64000\`, this would result in a VRAM usage of approximately \*\*2 GB\*\*. In this release, we introduce optional chunking over the batch dimension for the hidden states tensor during log-probability computation. This would cause the VRAM usage to be divided by the batch size or in this case be \*\*0.244 GB\*\*.This reduces the peak VRAM required to materialize hidden states, as reflected in the updated equation below: $$ \\text{Hidden States Memory (GB)} = \\frac{\\text{context length} \\times \\text{hidden states dim}}{1024^3} $$ Similar to our cross entropy loss in our \[500K Context Training\](/docs/blog/500k-context-length-fine-tuning.md) release, the new implementation \*\*automatically tunes hidden state batching\*\*. Users can also control this behavior via \`unsloth\_grpo\_mini\_batch\`. However, increasing \`unsloth\_grpo\_mini\_batch\` beyond the optimal value can introduce slight performance increase or slowdown (usually faster) compared to the previous loss function. However, during a GPT-OSS run (\`context\_length = 8192, batch\_size = 4, gradient\_accumulation\_steps = 2\`), setting \`unsloth\_grpo\_mini\_batch = 1\` and \`unsloth\_logit\_chunk\_multiplier = 4\` results in \*\*little to no speed degradation while reducing VRAM usage by approximately 5 GB\*\* compared to older versions of Unsloth. ![](https://unsloth.ai/files/LYOCU5JivEa9qakcugmk) {% hint style="success" %} \*\*Note:\*\* In Figures 3 and 4, we use the maximum effective batch size, which is 8 in this setup. The effective batch size is computed as \`batch\_size × gradient\_accumulation\_steps\`, giving \`4 × 2 = 8\`. For a deeper explanation of how effective batch sizes work in RL, see our \[advanced RL documentation\](/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation.md). {% endhint %} ### :cactus:Offloading activations for log softmax During the development of this release, we discovered that when tiling across the batch dimension for hidden states, the activations were not being offloaded after the fused logits and logprobs computation. Because logits are computed one batch at a time using \`hidden\_states\[i\] @ lm\_head\`, the existing activation offloading and gradient checkpointing logic, designed to operate within the model’s forward pass did not apply in this case. To address this, we added explicit logic to offload these activations outside the model’s forward pass, as shown in the Python pseudocode below: \`\`\`python class Unsloth\_Offloaded\_Log\_Softmax(torch.autograd.Function): def forward(...): with torch.no\_grad(): output = chunked\_hidden\_states\_selective\_log\_softmax(hidden\_states, lm\_head, ...) return output def backward(ctx, grad\_output): hidden\_states = ctx.saved\_hidden\_states hidden\_states.requires\_grad\_(True) with torch.enable\_grad(): output = chunked\_hidden\_states\_selective\_log\_softmax(hidden\_states, lm\_head, ...) torch.autograd.backward(output, grad\_output) return ... \`\`\` {% hint style="success" %} \*\*Note:\*\* This feature is only effective when chunking across the batch dimension or when \`unsloth\_grpo\_mini\_batch > 1\`. If all hidden states are materialized at once during the forward pass (i.e., \`unsloth\_grpo\_mini\_batch = 1\`), the backward pass requires the same amount of memory in the GPU regardless of whether activations are offloaded. Since activation offloading introduces a slight performance slowdown without reducing memory usage in this case, it provides no benefit. {% endhint %} ### :sparkles:Configuring parameters: If you do not configure \`unsloth\_grpo\_mini\_batch\` and \`unsloth\_logit\_chunk\_multiplier\`, we will \*\*automatically tune these two parameters\*\* for you based on your available VRAM and depending on the size of your context length. Below however is how you can change these variables in your GRPO run: \`\`\`python training\_args = GRPOConfig( ... unsloth\_grpo\_mini\_batch = 3 unsloth\_logit\_chunk\_multiplier = 2 ... ) \`\`\` A visualization of the optimizations and \`unsloth\_grpo\_mini\_batch\` and \`unsloth\_logit\_chunk\_multiplier\` can be seen in the diagram below. ![](https://unsloth.ai/files/v5qus5UAEf3474hNxoGP) The 3 matrices represent the overall larger batch or \`unsloth\_grpo\_mini\_batch\` (represented by the number of black brackets) and the rows of each of the matrices represents the context length that the \`unsloth\_logit\_chunk\_multiplier\` chunks the sequence length by (represented by the number of red brackets). ### :vhs:vLLM for RL \*\*For RL workflows, the inference/generation phase is the main bottleneck\*\*. To address this, we utilize \[vLLM\](https://github.com/vllm-project/vllm), which has accelerated generation by up to 11x compared to normal generation. Since GRPO was popularized last year, vLLM has been a core component of most RL frameworks including Unsloth. We want to extend our gratitude to the vLLM team and all its contributors for their work as they play a pivotal role in making Unsloth’s RL better! To try longer context RL, you can use any existing \[GRPO notebooks\](/docs/get-started/unsloth-notebooks.md#grpo-reasoning-rl-notebooks) (or update Unsloth if local): {% columns %} {% column width="33.33333333333333%" %} \[\*\*gpt-oss-20b\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) - GSPO {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[\*\*Qwen3-VL-8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) Vision RL {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[Qwen3-8B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_8B\_FP8\_GRPO.ipynb) L4 GPU {% embed url="" %} {% endcolumn %} {% endcolumns %} Acknowledgements: A huge thank you to the Hugging Face team and libraries for powering Unsloth and making this possible. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/nvfp4.md). # Run Unsloth Dynamic NVFP4 Guide Unsloth Dynamic NVFP4 is a quantized model format that runs on NVIDIA Blackwell GPUs and is designed for faster, more accurate 4-bit inference. It combines NVIDIA’s native NVFP4 precision with \[Unsloth Dynamic 2.0\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) quantization to preserve model accuracy while reducing VRAM usage and increasing speed. This guide explains FP4 quantization, compares NVFP4 with other formats, and shows how to run models like \[Gemma 4\](/docs/models/gemma-4.md) and \[Qwen3.6\](/docs/models/qwen3.6.md) locally using vLLM or SGLang on RTX 5050-5090, B200, RTX PRO 6000 and more GPUs. Dynamic NVFP4 works by selecting important layers to remain in FP8 (W8A8) or BF16 and the rest in W4A4 (not W4A16) instead of forcing every layer into FP4. This allows up to \*\*2.5x faster inference\*\* since W4A4 leverages Blackwell GPU's FP4 tensor cores. For all quants, we also provide FP8 KV cache calibration allowing for \*\*2x longer context lengths\*\*. {% hint style="success" %} \*\*All\*\* \[\*\*Gemma 4\*\*\](#gemma-4) \*\*models are now available as Unsloth Dynamic NVFP4 quants:\*\* E2B, E4B, 12B Unified, 26B-A4B MoE, and 31B Dense. Explore the \[Unsloth Dynamic NVFP4 Collection\](https://huggingface.co/collections/unsloth/nvfp4) for all our model uploads. {% endhint %} ### Float4 vs other precisions The trick for \*\*faster GPUs is to lower the numerical precision of matrix multiplications\*\*. The number of transistors needed for the matrix multiplication units is related to the \*\*square of the mantissa\*\*. The mantissa allows for numbers to have how many "fractional" decimals - so the more bits, the more accurate it can represent decimals. For example expressing 0.121332 is possible with more mantissa bits, whilst few mantissa bits will round it to 0.1. {% columns %} {% column width="50%" %} FP32 has 23 mantissa bits, so 23^2+ 8 exponent bits = 537 space is needed. Bfloat16 has 7 mantissa bits, so 7^2 + 8 exponent = 57 space. This means bfloat16 needs around 9x less space than FP32! And when we go to float8 which has 3 mantissa bits so 3^2 + 4 exponent = 13 - this is 41x less space than FP32! Finally float4 has 1 mantissa bit and 2 exponents so 3 space - a whopping 179x less space than FP32 - this essentially means a \*\*GPU can do around 179x more FP4 matrix multiplication than FP32 multiplication FLOPs in the same space\*\*! {% endcolumn %} {% column width="50%" %} !\[\](/files/glDYdAzUZzf4lOigOiNW) {% endcolumn %} {% endcolumns %} ### NVFP4 vs MXFP4 ![](https://unsloth.ai/files/bYvKn419Z2kmqDwoHJDW) ![](https://unsloth.ai/files/79gcz4UQlI97tZWt67tL) There is another FP4 format called MXFP4 - it's less accurate than NVFP4 due to 2 things: 1. NVFP4 uses a block size of 16 vs 32 for MXFP4 - this allows outliers to be isolated easier and scaling factors are provided for smaller subsets of weights which increases accuracy 2. A E4M3 (FP8) scale is used instead of a E8M0 (powers of 2 scaling) per block. Using a FP8 type block size looks to be much better especially for LLMs. ### Performance Analysis Our new dynamic NVFP4 Qwen3.6 quants run \\~\*\*2.5× faster\*\* than other NVFP4 quants, with \*\*better performance\*\* and comparable file sizes. Run Qwen3.6-27B NVFP4 \*\*2.5x faster\*\* on \*\*24GB VRAM\*\* and Qwen3.6-35B-A3B \*\*1.7x faster\*\* on \*\*32GB VRAM\*\*. We also added \*\*FP8 KV cache calibration\*\* for 2x longer context lengths! NVFP4 requires NVIDIA's Blackwell GPUs like RTX 50X, DGX Spark (see \[#dgx-spark-with-nvfp4-quants\](#dgx-spark-with-nvfp4-quants "mention")), B200, B300 GPUs. For older GPUs, our GGUFs work well! ![](https://unsloth.ai/files/6X0DPBb8ijDqd6cGk4sH) All benchmarks use 1x B200 128 concurrency. Higher concurrency can boost 35B to 17,561 tokens / s. We're also releasing two 35B-A3B NVFP4 versions: \* \[Qwen3.6-35B-A3B-NVFP4-Fast\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) which is a full W4A4 quant - 1.79x faster \* \[Qwen3.6-35B-A3B-NVFP4\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4) which is slightly bigger but more accurate and 1.56x faster For accuracy benchmarks, we conducted MMLU-Pro, AIME 2025, GPQA for FP8, BF16, NVIDIA's NVFP4 and our NVFP4s - we show our faster quants do similarly on all: ![](https://unsloth.ai/files/95OLTU6BPWYWYNmsJidJ) | Qwen3.6-35B-A3B | Qwen3.6-27B | | --- | --- | | [Qwen3.6-35B-A3B-NVFP4](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4)
(1.56x Faster) | [Qwen3.6-27B-NVFP4](https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4)
(2.5x Faster) | | [Qwen3.6-35B-A3B-NVFP4-Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast)
(1.79x Faster) | | \*\*MTP tensors are also built directly into the quants for additional speedups.\*\* Accuracy gains come from improvements to Qwen3.6’s chat template and dataset calibration. We use our previous chat template updates to help improve coding and tool-calling consistency while reducing looping and other reported issues. Our calibration uses a mix of our dataset optimized for coding, tool-calling and chat alongside UltraChat. For Decode speed (tokens per person), ours is 1.03x faster for 27B and 1.17x and 1.22x faster for 35B. ![](https://unsloth.ai/files/21SoRKrv07FpNP2b09iB) \### Overview Below are the hardware requirements for models which you can use including Gemma 4 and Qwen3.6. Also see the overall speed boost you will achieve: #### Gemma 4: | Gemma 4 variant | Required VRAM | Faster than BF16 | | ------------------------------------------------------------------ | ------------: | ---------------: | | \[E2B\](https://huggingface.co/unsloth/gemma-4-E2B-it-NVFP4) | 7 GB | 1.12× faster | | \[E4B\](https://huggingface.co/unsloth/gemma-4-E4B-it-NVFP4) | 9 GB | 1.22× faster | | \[12B Unified\](https://huggingface.co/unsloth/gemma-4-12b-it-NVFP4) | 11 GB | 1.26× faster | | \[26B A4B\](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-NVFP4) | 26 GB | 1.41× faster | | \[31B\](https://huggingface.co/unsloth/gemma-4-31B-it-NVFP4) | 32 GB | 1.45× faster | ![](https://unsloth.ai/files/MPXNED0IfRgdCqJFdUBt) \#### Qwen3.6: | Qwen3.6 variant | Required VRAM | Faster than other NVFP4 quants | | ------------------------------------------------------------------------- | ------------: | -----------------------------: | | \[27B\](https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4) | 24 GB | 2.5× faster | | \[35B A3B\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4) | 32 GB | 1.56× faster | | \[35B A3B Fast\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) | 32 GB | 1.79× faster | ### NVFP4 Benchmarks NVFP4 runs 4-bit weights and matrix multiplications directly on Blackwell Tensor Cores. Our Qwen3.6 NVFP4 quants use W4A4 so they actually use the FP4 tensor cores, so they decode faster than NVIDIA's which use W4A16. We also dynamically quantize layers to retain accuracy, and we conducted MMLU-Pro, AIME 2025, GPQA for all quants including comparing to FP8 and BF16. \*\*Qwen3.6-27B NVFP4 Accuracy Benchmarks\*\* | Provider | MMLU-Pro | GPQA | AIME 2025 | | -------- | -------: | ----: | --------: | | Unsloth | 86.25 | 86.34 | 93.12 | | NVIDIA | 85.96 | 86.87 | 93.12 | | FP8 | 86.11 | 86.87 | 93.75 | | BF16 | 85.96 | 88.13 | 93.33 | \*\*Qwen3.6-35B-A3B NVFP4 Accuracy Benchmarks\*\* | Provider | MMLU-Pro | GPQA | AIME 2025 | | ---------------- | -------: | ----: | --------: | | Unsloth | 85.85 | 86.74 | 92.29 | | \*\*Unsloth Fast\*\* | 85.58 | 87.75 | 91.67 | | NVIDIA | 85.60 | 87.12 | 91.88 | | FP8 | 85.75 | 86.74 | 93.12 | | BF16 | 85.75 | 86.36 | 92.50 | We also checked the output length of all benchmarks, and they are comparable, so the new NVFP4 quants do not think for longer which defeats the purpose of quantizing them! (Ie if it's 2x faster, but thinks 2x more, then that's useless) ![](https://unsloth.ai/files/5aW9XVFflv0cEGgWsgyJ) \## \*\*Run NVFP4 Tutorials\*\* To run NVFP4 quants, see below for commands to run Qwen3.6-27B in \[vLLM\](/docs/basics/inference-and-deployment/vllm-guide.md) and \[SGLang\](/docs/basics/inference-and-deployment/sglang-guide.md) (you can change model name to \`Qwen3.6-35-A3B-NVFP4\`). ### \*\*vLLM Tutorial\*\* You can run all NVFP4 models in \[vLLM\](https://github.com/vllm-project/vllm). Do NOT select any MoE backend - leave vLLM to select it - for eg Marlin is 2.5x slower! See \[#marlin-vs-flashinfer-vs-cutlass-vs-cute-dsl\](#marlin-vs-flashinfer-vs-cutlass-vs-cute-dsl "mention")If you have a DGX Spark, see \[#dgx-spark-serving\](#dgx-spark-serving "mention") you must use \`--moe-backend flashinfer\_b12x\` or you will get much slower inference. To install vLLM in a separate venv: {% code overflow="wrap" expandable="true" %} \`\`\`bash uv venv unsloth-nvfp4-env --python 3.13 source unsloth-nvfp4-env/bin/activate uv pip install "vllm>=0.25.0" "flashinfer-python>=0.6.13" "nvidia-cutlass-dsl>=4.5.2" \\ --torch-backend=auto \`\`\` {% endcode %} Then to serve the 35B Fast variant: \`\`\`shell vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast \`\`\` Change \`unsloth/Qwen3.6-35B-A3B-NVFP4-Fast\` to the NVFP4 quant names! To enable MTP / speculative decoding (faster decode but somewhat less throughput), use: \`\`\`bash vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast --speculative-config '{"method": "mtp", "num\_speculative\_tokens": 2}' \`\`\` If you get Torchcodec issues, be sure to do the below then relaunch vllm. {% code overflow="wrap" expandable="true" %} \`\`\`bash sudo apt-get update sudo apt-get install -y ffmpeg \`\`\` {% endcode %} ### \*\*DGX Spark Tutorial\*\* To ensure DGX Spark has the correct kernels (or you will get \*\*2x SLOWER inference\*\*), first check: {% code overflow="wrap" expandable="true" %} \`\`\`bash python -c " import torch; from vllm.utils.flashinfer import has\_flashinfer\_b12x\_gemm as g, has\_flashinfer\_b12x\_moe as m cap = torch.cuda.get\_device\_capability(); print('cap', cap, '| b12x gemm', g(), '| b12x moe', m()); assert cap\[0\] == 12 and g() and m(), 'b12x unavailable: serving would degrade to marlin W4A16'" \`\`\` {% endcode %} which should NOT error out - if it did, please update vllm or reinstall via: {% code overflow="wrap" expandable="true" %} \`\`\`bash uv venv unsloth-nvfp4-env --python 3.13 source unsloth-nvfp4-env/bin/activate uv pip install "vllm>=0.25.0" "flashinfer-python>=0.6.13" "nvidia-cutlass-dsl>=4.5.2" \\ --torch-backend=auto \`\`\` {% endcode %} Then to serve in vLLM for DGX Spark: {% code overflow="wrap" expandable="true" %} \`\`\`shellscript export CUTE\_DSL\_ARCH=sm\_121a vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast --moe-backend flashinfer\_b12x \`\`\` {% endcode %} If you get Torchcodec issues, be sure to do the below then relaunch vllm. {% code overflow="wrap" expandable="true" %} \`\`\`bash sudo apt-get update sudo apt-get install -y ffmpeg \`\`\` {% endcode %} ### \*\*SGLang Tutorial:\*\* You can run all NVFP4 models in \[SGLang\](https://github.com/sgl-project/sglang). Remember to switch out the model name for your desired model. \*\*Qwen3.6:\*\* \`\`\`bash python -m sglang.launch\_server --model-path unsloth/Qwen3.6-27B-NVFP4 --speculative-algorithm NEXTN \\ --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \`\`\` \*\*Gemma 4:\*\* \`\`\`bash python -m sglang.launch\_server --model-path unsloth/Gemma-4-31B-NVFP4 --speculative-algorithm NEXTN \\ --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \`\`\` ### Gemma 4 and others Every Gemma 4 variant now has an Unsloth Dynamic NVFP4 checkpoint. We show Gemma-4 having at most a 1.44x throughput boost on serving 128 concurrent people on 1x B200 vs BF16. Qwen3.5-122B-A10B is 1.38x faster and GLM-4.7-Flash is 1.27x faster. ![](https://unsloth.ai/files/dCihg7yYLunJQV9Lx1AK) \### Marlin vs Flashinfer vs cutlass vs cute-DSL We also found Marlin kernels to not support W4A4 well - enabling it will cause a 2.5x performance degradation - so use CUTLASS, Flashinfer-TRTLLM or Cute-DSL (auto enabled in vLLM)! Also if you have a DGX Spark, see \[#dgx-spark-serving\](#dgx-spark-serving "mention") you must use \`--moe-backend flashinfer\_b12x\` or you will get 2.5x slower inference. \*\*So don't set any backend - vLLM auto selects the best.\*\* | Model | scheme | backend | decode tok/s | thr out tok/s | | --------------- | ------ | ------------------- | ------------ | ------------- | | nvidia 27B | W4A16 | marlin (auto) | 115.6 | 2,403 | | unsloth 27B | W4A4 | marlin | 105.6 | 2,127 | | unsloth 27B | W4A4 | cutlass | 113.5 | 6,681 | | unsloth 27B | W4A4 | flashinfer\\\_trtllm | 112.6 | 6,158 | | unsloth 27B | W4A4 | \*\*cute-DSL (auto)\*\* | 125.9 | \*\*6,863\*\* | | nvidia 35B-A3B | W4A4 | marlin (auto) | 240.8 | 8,721 | | unsloth 35B-A3B | W4A4 | marlin | 215.8 | 8,619 | | unsloth 35B-A3B | W4A4 | cutlass | 158.3 | 11,017 | | unsloth 35B-A3B | W4A4 | \*\*cute-DSL (auto)\*\* | 295.2 | \*\*15,636\*\* | --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/nvfp4.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/kimi-k2.7-code.md). # Kimi K2.7 Code - How to Run Locally Kimi K2.7 Code is Moonshot AI’s agentic coding model, building on \[K2.6\](/docs/models/kimi-k2.6.md) to improve task completion while using \\~30% fewer thinking tokens. The 1T-parameter (32B active) MoE model supports thinking only, vision and 256K context. It delivers SOTA open performance across vision, coding, agentic, long-context, and chat tasks. Full precision requires 605GB of disk space; Unsloth \[Dynamic\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) 2-bit requires \*\*325GB (-48%)\*\*. Run \[\*\*Kimi-K2.7-Code-GGUF\*\*\](https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF) via Unsloth Studio or llama.cpp. \[\*\*Unsloth Dynamic\*\*\](/docs/basics/unsloth-dynamic-2.0-ggufs.md) \*\*quants\*\* upcasts important layers to 8-bit and 1-bit needs \*\*310GB+ VRAM/RAM\*\* setups\*\*.\*\* For \*\*lossless\*\* Kimi K2.6, use Q8 (\`UD-Q8\_K\_XL\`), which is only \*\*10GB larger\*\* than Q4 (\`UD-Q4\_K\_XL\`). You can run Kimi K2.7 Code via a Mac Studio or \[DGX Station\](/docs/blog/dgx-station.md). \*\*Table: Hardware requirements\*\* (units = total memory: RAM + VRAM, or unified memory) | Dynamic 1-bit | Dynamic 2-bit | Dynamic Q3 | Q8 (Lossless) | | ------------- | ------------- | ---------- | ------------- | | 310 GB | 325-350GB | 385-470 GB | 605 GB | ### 📊 Quantization Analysis Like \[Kimi-K2.6\](/docs/models/kimi-k2.6.md), \`UD-Q8\_K\_XL\` is lossless because Kimi uses int4 for MoE weights and BF16 for everything else, and \`Q8\_K\_XL\` follows that. Thus, we use the same Dynamic methodology for Kimi-K2.6 conversion. \`UD-Q4\_K\_XL\` is similar except the remaining tensors are \`Q8\_0\`, so it is near full precision and requires 600GB RAM/VRAM. \`UD-Q8\_K\_XL\` is 'truly lossless'. | Measurement | UD-Q2\\\_K\\\_XL | UD-Q4\\\_K\\\_XL | UD-Q8\\\_K\\\_XL (Lossless) | | ----------- | ------------ | ------------ | ----------------------- | | Disk Space | 339 GB | 584 GB | 595 GB | | Perplexity | \\~2.4131 | \\~1.8420 | \\~1.8419 | We followed \[jukofyork\](https://github.com/jukofyork)'s finding that \`const float d = max / -7;\` instead of the default \`const float d = max / -8;\` during the quantization process only on the MoE layers. This bijection patch on INT4-native MoEs allows the \`Q4\_0\` quant-type to reduce absolute error from 1.8% to near 0% (epsilon). For example below is the histogram for Kimi-K2.7-Code, and you can see -8 is unused entirely: ![](https://unsloth.ai/files/z2mIMc2S2DowK8M9U2rb) Note we must keep other layers in BF16 as well and not smart "Q4\\\_0". We show below the error plots for both versus the BF16 baseline. \`UD-Q8-K\_XL\` is truly "lossless" with some machine epsilon difference when converting Q4\\\_0 to BF16. So Q4\\\_K\\\_XL does have some quantization error due to Q8\\\_0 being used, whilst Q8\\\_K\\\_XL is nearly lossless, except for BF16 rounding. ![](https://unsloth.ai/files/qO1cObp9kH6T23i9bWjK) For Q4\\\_K\\\_XL, we also plot the per tensor error from Q8\\\_0 vs BF16 as well. In general there is some error between Q8\\\_K\\\_XL (near lossless) vs Q4\\\_K\\\_XL, but not much. ![](https://unsloth.ai/files/SgUUfEfdLpJyKmrs3cK0) \### :gear: Usage Guide Kimi K2.7 Code is \*\*thinking-only\*\*, with \*\*\`preserve\_thinking\` always enabled\*\*. Instant mode is not supported. | Default (Thinking Mode) | | ----------------------- | | temperature = 1.0 | | top\\\_p = 0.95 | \* Suggested context length = \`98,304\` (up to \`262,144\`) If the model fits, you will get >100 tokens/s when using B200s. We recommend \`UD-Q2\_K\_XL\` (345GB) as a good size/quality balance. Best rule of thumb: RAM+VRAM ≈ the quant size; otherwise it’ll still work, just slower due to offloading. #### Chat Template for Kimi K2.7-Code Running \`tokenizer.apply\_chat\_template(\[{"role": "user", "content": "What is 1+1?"},\])\` gets: {% code overflow="wrap" %} \`\`\` <|im\_user|>user<|im\_middle|>What is 1+1?<|im\_end|><|im\_assistant|>assistant<|im\_middle|> \`\`\` {% endcode %} If we also input tools as referenced in \[Tool Calling Guide\](/docs/basics/tool-calling-guide-for-local-llms.md), then we see the below: {% code overflow="wrap" expandable="true" %} \`\`\` <|im\_system|>tool\_declare<|im\_middle|># Tools ## functions namespace functions { // Add two numbers. type add\_number = (\_: { // The first number. a: string, // The second number. b: string }) => any; // Multiply two numbers. type multiply\_number = (\_: { // The first number. a: string, // The second number. b: string }) => any; // Subtract two numbers. type subtract\_number = (\_: { // The first number. a: string, // The second number. b: string }) => any; // Writes a random story. type write\_a\_story = (\_: {}) => any; // Perform operations from the terminal. type terminal = (\_: { // The command you wish to launch, e.g \`ls\`, \`rm\`, ... command: string }) => any; // Call a Python interpreter with some Python code that will be ran. type python = (\_: { // The Python code to run code: string }) => any; } <|im\_end|><|im\_user|>user<|im\_middle|>What is 1+1?<|im\_end|><|im\_assistant|>assistant<|im\_middle|> \`\`\` {% endcode %} ## Run Kimi K2.7 Code Guide ### 🦥 Run Kimi-K2.7-Code in Unsloth Studio Kimi K2.7 Code can run in \[Unsloth Studio\](/docs/new/studio.md), an open-source web UI for local AI. \*\*Unsloth Studio automatically offloads to RAM and detects multiGPU setups\*\*. With Unsloth Studio, you can run models locally on \*\*MacOS, Windows\*\*, Linux and: {% columns %} {% column %} \* Search, download, \[run GGUFs\](/docs/new/studio.md#run-models-locally) and safetensor models \* \[\*\*Self-healing\*\* tool calling\](/docs/new/studio.md#execute-code--heal-tool-calling) + \*\*web search\*\* \* \[\*\*Code execution\*\*\](/docs/new/studio.md#run-models-locally) (Python, Bash) \* \[Automatic inference\](/docs/new/studio.md#model-arena) parameter tuning (temp, top-p, etc.) \* Fast CPU + GPU inference via llama.cpp \* \[Train LLMs\](/docs/new/studio.md#no-code-training) 2x faster with 70% less VRAM {% endcolumn %} {% column %} ![](https://unsloth.ai/files/dQy5izI8WRumFBHqVXtW) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} \*\*Install and Launch Unsloth\*\* To install, run in your terminal: MacOS, Linux, WSL: \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Windows PowerShell: \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` \*\*Launch Unsloth\*\* MacOS, Linux, WSL and Windows: \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` Then open \`http://127.0.0.1:8888\` (or your specific URL) in your browser. {% endstep %} {% step %} \*\*Search and download Kimi K2.7-Code\*\* Unsloth Studio automatically offloads to RAM and detects multiGPU setups. On first launch you will need to create a password to secure your account and sign in again later. Then go to the \[Unsloth Chat\](/docs/new/studio/chat.md) tab and search for \*\*Kimi-K2.7 Code\*\* in the search bar and download your desired model and quant. Ensure you have enough compute the run the model. ![](https://unsloth.ai/files/ppkx2t5PvxZXVCFlyZPH) {% endstep %} {% step %} \*\*Run Kimi-K2.7-Code\*\* Inference parameters should be auto-set when using Unsloth Studio, however you can still change it manually. You can also edit the context length, chat template and other settings. For more information, you can view our \[Unsloth Studio inference guide\](/docs/new/studio/chat.md). ![](https://unsloth.ai/files/PxQ3x37GwzkPPjHW6pVh) Example of Qwen3.6 running with tool-calling {% endstep %} {% endstepper %} ### 🦙 Run Kimi K2.7 Code in llama.cpp For this guide we'll be running the \`UD-Q2\_K\_XL\` quant which will require at least 345GB RAM. Feel free to change quantization type. GGUF: \[\*\*Kimi-K2.7-Code-GGUF\*\*\](https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF) For these tutorials, we will using \[llama.cpp\](llama.cpphttps://github.com/ggml-org/llama.cpp) for fast local inference, especially if you have a CPU. {% stepper %} {% step %} Obtain the latest \`llama.cpp\` \*\*on\*\* \[\*\*GitHub here\*\*\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} \*\*Let's first get an image!\*\* You can also upload images as well. We shall use , which is just our mini logo showing how finetunes are made with Unsloth: {% code overflow="wrap" %} \`\`\`bash wget https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/unsloth%20made%20with%20love.png -O unsloth.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/6grTWZGCnR0olP2Q8SgB) Let's get the 2nd image at {% code overflow="wrap" %} \`\`\`bash wget https://files.worldwildlife.org/wwfcmsprod/images/Sloth\_Sitting\_iStock\_3\_12\_2014/story\_full\_width/8l7pbjmj29\_iStock\_000011145477Large\_mini\_\_1\_.jpg -O picture.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/Bo2dlyVZZXbxGr2zhpKU) {% endstep %} {% step %} You can now use \`llama.cpp\` directly to load and download models, just like \`ollama run\`. First, select the quantization type you want like \`Q2\_K\_XL\`. Also use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. Note this download process might be very slow, so it's probably best to use the manual download process in the next section. \`\`\`bash export LLAMA\_CACHE="unsloth/Kimi-K2.7-Code-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/Kimi-K2.7-Code-GGUF:UD-Q2\_K\_XL \\ --temp 1.0 \\ --top-p 0.95 \`\`\` {% endstep %} {% step %} If you want to download the model manually, we can download the model via the code below (after installing \`pip install huggingface\_hub\`). If downloads get stuck, see: \[Hugging Face Hub, XET debugging\](/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) \`\`\`bash hf download unsloth/Kimi-K2.7-Code-GGUF \\ --local-dir unsloth/Kimi-K2.7-Code-GGUF \\ --include "\*mmproj-F16\*" \\ --include "\*UD-Q2\_K\_XL\*" # Use "\*UD-Q8\_K\_XL\*" for full precision \`\`\` {% endstep %} {% step %} Then run the model in conversation mode: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Kimi-K2.7-Code-GGUF/UD-Q2\_K\_XL/Kimi-K2.7-Code-UD-Q2\_K\_XL-00001-of-00008.gguf \\ --mmproj unsloth/Kimi-K2.7-Code-GGUF/mmproj-F16.gguf \\ --temp 1.0 \\ --top-p 0.95 \`\`\` {% endcode %} Then you will see the below:\\ !\[\](/files/ob2iB0C92LI2catVxe3a) {% endstep %} {% step %} Then use \`/image\` to load both images in and ask "What is this image": ![](https://unsloth.ai/files/yYioexUrNFFhr0wuT05X) and you will get something like below: ![](https://unsloth.ai/files/UNZ5UG45F602UK4L3tHi) On the 2nd image of the sloth: ![](https://unsloth.ai/files/YuSfYCovocMIoyYXZaJU) Which will get you: ![](https://unsloth.ai/files/qR6N1jloqjYHJNM87h6V) {% endstep %} {% endstepper %} ### 📊 Benchmarks You can view further below for benchmarks in table format: ![](https://unsloth.ai/files/0FziWILjKIV7UpB3HZPl) | Benchmark | Kimi K2.7 Code | Kimi K2.6 | GPT-5.5 | Claude Opus 4.8 | | :------------------: | :------------: | :-------: | :-----: | :-------------: | | \*\*Coding\*\* | | | | | | Kimi Code Bench v2 | 62.0 | 50.9 | 69.0 | 67.4 | | Program Bench | 53.6 | 48.3 | 69.1 | 63.8 | | MLS Bench Lite | 35.1 | 26.7 | 35.5 | 42.8 | | \*\*Agentic\*\* | | | | | | Kimi Claw 24/7 Bench | 46.9 | 42.9 | 52.8 | 50.4 | | MCP Atlas | 76.0 | 69.4 | 79.4 | 81.3 | | MCP Mark Verified | 81.1 | 72.8 | 92.9 | 76.4 | --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/kimi-k2.7-code.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/preference-dpo-orpo-and-kto.md). # Entraînement d'optimisation des préférences - DPO, ORPO et KTO DPO (Direct Preference Optimization), ORPO (Odds Ratio Preference Optimization), PPO, KTO Reward Modelling fonctionnent tous avec Unsloth. Nous disposons de notebooks Google Colab pour reproduire GRPO, ORPO, DPO Zephyr, KTO et SimPO : \* \[Notebooks GRPO\](/docs/fr/commencer/unsloth-notebooks.md#grpo-reasoning-rl-notebooks) \* \[Notebook ORPO\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-ORPO.ipynb) \* \[Notebook DPO Zephyr\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Zephyr\_\\(7B\\)-DPO.ipynb) \* \[Notebook KTO\](https://colab.research.google.com/drive/1MRgGtLWuZX4ypSfGguFgC-IblTvO2ivM?usp=sharing) \* \[Notebook SimPO\](https://colab.research.google.com/drive/1Hs5oQDovOay4mFA6Y9lQhVJ8TnbFLFh2?usp=sharing) Nous sommes également dans la documentation officielle de 🤗Hugging Face ! Nous sommes dans le \[docs SFT\](https://huggingface.co/docs/trl/main/en/sft\_trainer#accelerate-fine-tuning-2x-using-unsloth) et le \[docs DPO\](https://huggingface.co/docs/trl/main/en/dpo\_trainer#accelerate-dpo-fine-tuning-using-unsloth). ## Code DPO \`\`\`python import os os.environ\["CUDA\_VISIBLE\_DEVICES"\] = "0" # Optionnel : définir l'ID du dispositif GPU from unsloth import FastLanguageModel, PatchDPOTrainer from unsloth import is\_bfloat16\_supported PatchDPOTrainer() import torch from trl import DPOTrainer, DPOConfig # Changé depuis TrainingArguments model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/zephyr-sft-bnb-4bit", max\_seq\_length = max\_seq\_length, dtype = None, load\_in\_4bit = True, ) # Effectuer le patching du modèle et ajouter des poids LoRA rapides model = FastLanguageModel.get\_peft\_model( model, r = 64, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\], lora\_alpha = 64, lora\_dropout = 0, # Prend en charge n'importe quelle valeur, mais = 0 est optimisé bias = "none", # Prend en charge n'importe quelle valeur, mais = "none" est optimisé # \[NOUVEAU\] "unsloth" utilise 30 % de VRAM en moins, permet des tailles de batch 2x plus grandes ! use\_gradient\_checkpointing = "unsloth", # True ou "unsloth" pour des contextes très longs random\_state = 3407, max\_seq\_length = max\_seq\_length, ) dpo\_trainer = DPOTrainer( model = model, ref\_model = None, args = DPOConfig( # Utiliser DPOConfig per\_device\_train\_batch\_size = 4, gradient\_accumulation\_steps = 8, warmup\_ratio = 0.1, num\_train\_epochs = 3, fp16 = not is\_bfloat16\_supported(), bf16 = is\_bfloat16\_supported(), logging\_steps = 1, optim = "adamw\_8bit", seed = 42, output\_dir = "outputs", ), beta = 0.1, train\_dataset = YOUR\_DATASET\_HERE, # eval\_dataset = YOUR\_DATASET\_HERE, tokenizer = tokenizer, max\_length = 1024, max\_prompt\_length = 512, ) dpo\_trainer.train() \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/preference-dpo-orpo-and-kto.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/intel.md). # Fine-tuning LLMs on Intel GPUs with Unsloth You can now fine-tune LLMs on your local Intel device with Unsloth! Read our guide on exactly how to get started with training your own custom model. Before you begin, make sure you have: \* \*\*Intel GPU:\*\* Data Center GPU Max Series, Arc Series, or Intel Ultra AIPC \* \*\*OS:\*\* Linux (Ubuntu 22.04+ recommended) or Windows 11 (recommended) \* \*\*Windows only:\*\* Install Intel oneAPI Base Toolkit 2025.2.1 (select version 2025.2.1) \* \*\*Intel Graphics driver:\*\* Latest recommended driver for Windows/Linux \* \*\*Python:\*\* 3.10+ ### Build Unsloth with Intel Support {% stepper %} {% step %} #### Create a new conda environment (Optional) \`\`\`bash conda create -n unsloth-xpu python==3.10 conda activate unsloth-xpu \`\`\` {% endstep %} {% step %} #### Install Unsloth \`\`\`bash git clone https://github.com/unslothai/unsloth.git cd unsloth pip install .\[intel-gpu-torch290\] \`\`\` {% hint style="info" %} Linux Only: Install \[vLLM\](/docs/basics/inference-and-deployment/vllm-guide.md) (Optional)\\ You can also install vLLM for \[inference\](/docs/basics/inference-and-deployment.md) and \[RL\](/docs/get-started/reinforcement-learning-rl-guide.md). Please follow \[vLLM's guide\](https://docs.vllm.ai/en/latest/getting\_started/installation/gpu/#intel-xpu). {% endhint %} {% endstep %} {% step %} #### Verify your environments \`\`\`python import torch print(f"PyTorch version: {torch.\_\_version\_\_}") print(f"XPU available: {torch.xpu.is\_available()}") print(f"XPU device count: {torch.xpu.device\_count()}") print(f"XPU device name: {torch.xpu.get\_device\_name(0)}") \`\`\` {% endstep %} {% step %} #### Start fine-tuning. You can directly use our Unsloth \[notebooks\](/docs/get-started/unsloth-notebooks.md) or view our dedicated \[fine-tuning\](/docs/get-started/fine-tuning-llms-guide.md) or \[reinforcement learning\](/docs/get-started/reinforcement-learning-rl-guide.md) guides. {% endstep %} {% endstepper %} ### Windows Only - Runtime Configurations In Command Prompt with Administrator privilege, enable long path support in the Windows registry: \`\`\`bash powershell -Command "Set-ItemProperty -Path "HKLM:\\\\SYSTEM\\\\CurrentControlSet\\\\Control\\\\FileSystem" -Name "LongPathsEnabled" -Value 1 \`\`\` This command only needs to be set once on a single machine. It does not need to be configured before each run. Then: 1. Download level-zero-win-sdk-1.20.2.zip from \[GitHub\](https://github.com/oneapi-src/level-zero/releases/tag/v1.20.2) 2. Unzip the level-zero-win-sdk-1.20.2.zip 3. In Command Prompt, under conda environment unsloth-xpu: \`\`\`bash call "C:\\Program Files (x86)\\Intel\\oneAPI\\setvars.bat" - set ZE\_PATH=path\\to\\the\\unzipped\\level-zero-win-sdk-1.20.2 \`\`\` ### Example 1: QLoRA Fine-tuning with SFT This example demonstrates how to fine-tune a Qwen3-32B model using 4-bit QLoRA on an Intel GPU. QLoRA significantly reduces memory requirements, making it possible to fine-tune large models on consumer-grade hardware. {% code expandable="true" %} \`\`\`python from unsloth import FastLanguageModel, FastModel from trl import SFTTrainer, SFTConfig from datasets import load\_dataset max\_seq\_length = 2048 # Supports RoPE Scaling internally, so choose any! # Get LAION dataset url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified\_chip2.jsonl" dataset = load\_dataset("json", data\_files = {"train" : url}, split = "train") # 4bit pre quantized models we support for fast downloading + no OOMs. fourbit\_models = \[ "unsloth/Qwen3-32B-bnb-4bit", "unsloth/Qwen3-14B-bnb-4bit", "unsloth/Qwen3-8B-bnb-4bit", "unsloth/Qwen3-4B-bnb-4bit", "unsloth/Qwen3-1.7B-bnb-4bit", "unsloth/Qwen3-0.6B-bnb-4bit", # "unsloth/Qwen2.5-32B-bnb-4bit", # "unsloth/Qwen2.5-14B-bnb-4bit", # "unsloth/Qwen2.5-7B-bnb-4bit", # "unsloth/Qwen2.5-3B-bnb-4bit", # "unsloth/Qwen2.5-1.5B-bnb-4bit", # "unsloth/Qwen2.5-0.5B-bnb-4bit", # "unsloth/Llama-3.2-3B-bnb-4bit", # "unsloth/Llama-3.2-1B-bnb-4bit", # "unsloth/Llama-3.1-8B-bnb-4bit", # "unsloth/Llama-3.1-70B-bnb-4bit", # "unsloth/mistral-7b-bnb-4bit", # "unsloth/Phi-4", # "unsloth/Phi-3.5-mini-instruct", # "unsloth/Phi-3-medium-4k-instruct", # "unsloth/Phi-3-mini-4k-instruct", # "unsloth/gemma-2-9b-bnb-4bit", # "unsloth/gemma-2-27b-bnb-4bit", \] # More models at https://huggingface.co/unsloth model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/Qwen3-32B-bnb-4bit", max\_seq\_length = max\_seq\_length, load\_in\_4bit = True, # token = "hf\_...", # use one if using gated models like meta-llama/Llama-2-7b-hf ) model = FastLanguageModel.get\_peft\_model( model, r = 16, # Choose any number > 0 ! Suggested 8, 16, 32, 64, 128 target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj", \], lora\_alpha = 16, lora\_dropout = 0, # Supports any, but = 0 is optimized bias = "none", # Supports any, but = "none" is optimized use\_gradient\_checkpointing = "unsloth", # True or "unsloth" for very long context random\_state = 3407, use\_rslora = False, # We support rank stabilized LoRA loftq\_config = None, # And LoftQ ) trainer = SFTTrainer( model = model, tokenizer = tokenizer, train\_dataset = dataset, dataset\_text\_field = "text", max\_seq\_length = max\_seq\_length, dataset\_num\_proc = 1, # Recommended on Windows packing = False, # Can make training 5x faster for short sequences. args = SFTConfig( per\_device\_train\_batch\_size = 2, gradient\_accumulation\_steps = 4, warmup\_steps = 5, max\_steps = 60, learning\_rate = 2e-4, logging\_steps = 1, optim = "adamw\_8bit", weight\_decay = 0.01, lr\_scheduler\_type = "linear", seed = 3407, dataset\_num\_proc=1, # Recommended on Windows ), ) trainer.train() \`\`\` {% endcode %} ### Example 2: Reinforcement Learning GRPO GRPO is a \[reinforcement learning\](/docs/get-started/reinforcement-learning-rl-guide.md) technique for aligning language models with human preferences. This example shows how to train a model to follow a specific XML output format using multiple reward functions. #### What is GRPO? GRPO improves upon traditional RLHF by: \* Using group-based normalization for more stable training \* Supporting multiple reward functions for multi-objective optimization \* Being more memory efficient than PPO {% code expandable="true" %} \`\`\`python from unsloth import FastLanguageModel import re from trl import GRPOConfig, GRPOTrainer from datasets import load\_dataset, Dataset max\_seq\_length = 1024 # Can increase for longer reasoning traces lora\_rank = 32 # Larger rank = smarter, but slower max\_prompt\_length = 256 # Load and prep dataset SYSTEM\_PROMPT = """ Respond in the following format: ... ... """ XML\_COT\_FORMAT = """\\ {reasoning} {answer} """ def extract\_xml\_answer(text: str) -> str: answer = text.split("")\[-1\] answer = answer.split("")\[0\] return answer.strip() def extract\_hash\_answer(text: str) -> str | None: if "####" not in text: return None return text.split("####")\[1\].strip() # uncomment middle messages for 1-shot prompting def get\_gsm8k\_questions(split: str = "train") -> Dataset: data = load\_dataset("openai/gsm8k", "main")\[split\] # type: ignore data = data.map( lambda x: { # type: ignore "prompt": \[ {"role": "system", "content": SYSTEM\_PROMPT}, {"role": "user", "content": x\["question"\]}, \], "answer": extract\_hash\_answer(x\["answer"\]), } ) # type: ignore return data # type: ignore # Reward functions def correctness\_reward\_func(prompts, completions, answer, \*\*kwargs) -> list\[float\]: responses = \[completion\[0\]\["content"\] for completion in completions\] q = prompts\[0\]\[-1\]\["content"\] extracted\_responses = \[extract\_xml\_answer(r) for r in responses\] print( "-" \* 20, f"Question:\\n{q}", f"\\nAnswer:\\n{answer\[0\]}", f"\\nResponse:\\n{responses\[0\]}", f"\\nExtracted:\\n{extracted\_responses\[0\]}", ) return \[2.0 if r == a else 0.0 for r, a in zip(extracted\_responses, answer)\] def int\_reward\_func(completions, \*\*kwargs) -> list\[float\]: responses = \[completion\[0\]\["content"\] for completion in completions\] extracted\_responses = \[extract\_xml\_answer(r) for r in responses\] return \[0.5 if r.isdigit() else 0.0 for r in extracted\_responses\] def strict\_format\_reward\_func(completions, \*\*kwargs) -> list\[float\]: """Reward function that checks if the completion has a specific format.""" pattern = r"^\\n.\*?\\n\\n\\n.\*?\\n\\n$" responses = \[completion\[0\]\["content"\] for completion in completions\] matches = \[re.match(pattern, r) for r in responses\] return \[0.5 if match else 0.0 for match in matches\] def soft\_format\_reward\_func(completions, \*\*kwargs) -> list\[float\]: """Reward function that checks if the completion has a specific format.""" pattern = r".\*?\\s\*.\*?" responses = \[completion\[0\]\["content"\] for completion in completions\] matches = \[re.match(pattern, r) for r in responses\] return \[0.5 if match else 0.0 for match in matches\] def count\_xml(text: str) -> float: count = 0.0 if text.count("\\n") == 1: count += 0.125 if text.count("\\n\\n") == 1: count += 0.125 if text.count("\\n\\n") == 1: count += 0.125 count -= len(text.split("\\n\\n")\[-1\]) \* 0.001 if text.count("\\n") == 1: count += 0.125 count -= (len(text.split("\\n")\[-1\]) - 1) \* 0.001 return count def xmlcount\_reward\_func(completions, \*\*kwargs) -> list\[float\]: contents = \[completion\[0\]\["content"\] for completion in completions\] return \[count\_xml(c) for c in contents\] if \_\_name\_\_ == "\_\_main\_\_": model, tokenizer = FastLanguageModel.from\_pretrained( model\_name="unsloth/Qwen3-0.6B", max\_seq\_length=max\_seq\_length, load\_in\_4bit=False, # False for LoRA 16bit fast\_inference=False, # Enable vLLM fast inference max\_lora\_rank=lora\_rank, gpu\_memory\_utilization=0.7, # Reduce if out of memory device\_map="xpu:0", ) model = FastLanguageModel.get\_peft\_model( model, r=lora\_rank, # Choose any number > 0 ! Suggested 8, 16, 32, 64, 128 target\_modules=\[ "q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj", \], # Remove QKVO if out of memory lora\_alpha=lora\_rank, use\_gradient\_checkpointing="unsloth", # Enable long context finetuning random\_state=3407, ) dataset = get\_gsm8k\_questions() training\_args = GRPOConfig( learning\_rate=5e-6, adam\_beta1=0.9, adam\_beta2=0.99, weight\_decay=0.1, warmup\_ratio=0.1, lr\_scheduler\_type="cosine", optim="adamw\_torch", logging\_steps=1, per\_device\_train\_batch\_size=1, gradient\_accumulation\_steps=1, # Increase to 4 for smoother training num\_generations=4, # Decrease if out of memory max\_prompt\_length=max\_prompt\_length, max\_completion\_length=max\_seq\_length - max\_prompt\_length, # num\_train\_epochs=1, # Set to 1 for a full training run max\_steps=20, save\_steps=250, max\_grad\_norm=0.1, report\_to="none", # Can use Weights & Biases output\_dir="outputs", ) trainer = GRPOTrainer( model=model, processing\_class=tokenizer, reward\_funcs=\[ xmlcount\_reward\_func, soft\_format\_reward\_func, strict\_format\_reward\_func, int\_reward\_func, correctness\_reward\_func, \], args=training\_args, train\_dataset=dataset, dataset\_num\_proc=1, # Recommended on Windows ) trainer.train() \`\`\` {% endcode %} ## Troubleshooting ### Out of Memory (OOM) Errors If you run out of memory, try these solutions: 1. \*\*Reduce batch size:\*\* Lower \`per\_device\_train\_batch\_size\`. 2. \*\*Use a smaller model:\*\* Start with a smaller model to reduce memory requirements. 3. \*\*Reduce sequence length:\*\* Lower \`max\_seq\_length\`. 4. \*\*Reduce LoRA rank:\*\* Use \`r=8\` instead of \`r=16\` or \`r=32\`. 5. \*\*For GRPO, reduce number of generations:\*\* Lower \`num\_generations\`. ### (Windows Only) Intel Ultra AIPC iGPU Shared Memory For Intel Ultra AIPC with recent GPU drivers on Windows, the shared GPU memory for the integrated GPU typically defaults to \*\*57%\*\* of system memory. For larger models (e.g., \*\*Qwen3-32B\*\*), or when using longer max sequence length, larger batch size, LoRA adapters with larger LoRA rank, etc., during fine-tuning, you could increase available VRAM by raising the percentage of system memory allocated to the iGPU. You can adjust this by modifying the registry: \* Path: \`Computer\\HKEY\_LOCAL\_MACHINE\\SYSTEM\\CurrentControlSet\\Control\\GraphicsDrivers\\MemoryManager\` \* Key to change:\\ \`SystemPartitionCommitLimitPercentage\` (set to a larger percentage) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/intel.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connect-curl-and-http-to-unsloth.md). # Connect Curl & HTTP to Unsloth Unsloth exposes three OpenAI/Anthropic-compatible wire formats at the same base URL on the port Unsloth started on. All of them take an \`Authorization: Bearer sk-unsloth-…\` header and return either JSON or SSE, depending on whether you set \`stream\`. \\ \\ This page groups the recipes by endpoint (\`/v1/chat/completions\`, \`/v1/messages\`, \`/v1/responses\`, \`/v1/models\`) and ends with a shared section on Unsloth's built-in \*\*server-side tools\*\*, which work across all the chat endpoints. {% hint style="info" %} If you're not sure what URL / key / model name to use, read the API overview first. It walks you through starting Unsloth, loading a model, and creating an \`sk-unsloth-…\` key. {% endhint %} ### 🔑 Authentication Every request needs an \`Authorization\` header: \`\`\` Authorization: Bearer sk-unsloth-xxxxxxxxxxxx \`\`\` To keep keys out of your shell history, export the key once and reference the env var: \`\`\`bash export UNSLOTH\_STUDIO\_AUTH\_TOKEN=sk-unsloth-xxxxxxxxxxxx \`\`\` The snippets below inline the key as \`sk-unsloth-xxxxxxxxxxxx\` for clarity. In practice, substitute \`$UNSLOTH\_STUDIO\_AUTH\_TOKEN\`. ### 📋 List loaded models \`\`\`bash curl http://localhost:8888/v1/models \\ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \`\`\` Response: \`\`\`json { "object": "list", "data": \[ {"id": "unsloth/gemma-3-27b-it-GGUF", "object": "model", "owned\_by": "local"} \] } \`\`\` ![](https://unsloth.ai/files/P7uayFTXkv36B22HRBye) Use the \`id\` field whenever a request needs a \`"model"\` value (or when a client like opencode asks for a \*\*Model ID\*\*). ### 💬 Chat Completions (\`/v1/chat/completions\`) The OpenAI Chat Completions dialect. The broadest compatibility surface. Works with the OpenAI SDK, opencode, Cursor, Continue, Cline, Open WebUI, SillyTavern, and most OpenAI-compatible tools. #### Basic request \`\`\`bash curl http://localhost:8888/v1/chat/completions \\ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \\ -H "Content-Type: application/json" \\ -d '{ "model": "default", "messages": \[{"role": "user", "content": "Hello"}\] }' \`\`\` ![](https://unsloth.ai/files/tXGzWlxxh76cnIYXtaxX) \#### Streaming Add \`"stream": true\` and the response switches to Server-Sent Events (\`text/event-stream\`). Tell \`curl\` to flush as bytes arrive with \`--no-buffer\` (\`-N\`): \`\`\`bash curl -N http://localhost:8888/v1/chat/completions \\ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \\ -H "Content-Type: application/json" \\ -d '{ "model": "qwen-local", "messages": \[{"role": "user", "content": "Write a haiku about locally-run LLMs."}\], "stream": true }' \`\`\` Each line of the response looks like \`data: {"choices":\[{"delta":{"content":"..."}}\]}\`, ending with \`data: \[DONE\]\`. ![](https://unsloth.ai/files/nrGfv5aISlkPONDyOlJP) \#### Images (vision) Attach an image as an \`image\_url\` content part in the user message. The URL can be HTTPS or a base64 \`data:\` URI: \`\`\`bash # Embed a local file as base64 (trimmed for brevity) IMG=$(base64 -w 0 test.jpg) cat > /tmp/request.json /tmp/request.json For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/mistral-3.5.md). # Mistral 3.5 - How To Run Locally Mistral releases Mistral-Medium-3.5-128B, their new dense 128B parameter, multimodal, hybrid reasoning model. It supports text and image input, text output, a 256K context window and excels at reasoning, coding, long-context, tool use, agentic workflows, and multimodal doc/image understanding. Mistral Medium 3.5 offers highly competitive performance for models 5x its size. Run locally on \\~64GB RAM. GGUF: \[Mistral-Medium-3.5-128B-GGUF\](https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF) {% hint style="success" %} \*\*May 1, 2026 Update:\*\* We worked with Mistral to fix Mistral Medium 3.5 inference affecting some implementations, and released updated GGUFs with the fix (\*\*NOT related to Unsloth\*\* or our quants). The issue was caused by a YaRN parsing quirk affecting several implementations, including \`transformers\` and \`llama.cpp\`. Changing \`mscale\_all\_dim\` from \`1\` to \`0\` resolved it. We also fixed \`mmproj\` files not being generated correctly. \*\*Mistral has now pushed our fixes to their official repo!\*\* {% endhint %} ### Usage Guide {% hint style="info" %} Vision for GGUFs it now supported for now. Support will come later. {% endhint %} Table: Mistral Medium 3.5 recommended hardware requirements. Units are total memory: RAM + VRAM, or unified memory. | Mistral 3.5 | 3-bit | 4-bit | 8-bit | | --------------- | ----- | ----- | ---------- | | Medium 3.5 128B | 64 GB | 80 GB | 128-170 GB | {% hint style="info" %} Your total available memory should at least exceed the size of the quantized model you download. If it does not, llama.cpp can still run with partial RAM / disk offload, but generation will be slower. You will also need more memory for long context, larger batches, tool-heavy agent runs and image prompts. {% endhint %} #### Recommended Settings Use Mistral's recommended reasoning settings: \* \`reasoning\_effort="none"\` → fast instant replies, chat, extraction and simple instructions. \* \`reasoning\_effort="high"\` → reasoning mode, recommended for complex prompts, coding, research, math and agentic usage. Recommended sampling defaults: \* Use \`temperature = 0.7\` for \`reasoning\_effort="high"\`. \* Use \`temperature = 0.0\` to \`0.7\` for \`reasoning\_effort="none"\`, depending on the task. \* Keep repetition and presence penalties disabled or at \`1.0\` unless you see looping. \* Maximum context length of \`262,144\` #### \*\*Reasoning Mode\*\* Mistral Medium 3.5 supports instant instruct mode and reasoning mode with a 'high' option. To enable high reasoning for llama.cpp / llama-server: \`\`\`bash --chat-template-kwargs '{"reasoning\_effort":"high"}' \`\`\` To disable reasoning: \`\`\`bash --chat-template-kwargs '{"reasoning\_effort":"none"}' \`\`\` If you're on Windows PowerShell, use: \`\`\`powershell --chat-template-kwargs "{\\"reasoning\_effort\\":\\"none\\"}" \`\`\` ## Run Mistral 3.5 Tutorials Because Mistral Medium 3.5 is a dense 128B model, the recommended starting point is Dynamic 4-bit GGUFs for local inference. GGUF: \`unsloth/Mistral-Medium-3.5-128B-GGUF\` [Run in Unsloth Studio](https://unsloth.ai/pages/1jdqbevmbwD2ZUqVORgk#unsloth-studio-guide) [Run in llama.cpp](https://unsloth.ai/pages/1jdqbevmbwD2ZUqVORgk#llama.cpp-guide) {% hint style="warning" %} Currently no multimodal/vision GGUF works in \*\*Ollama\*\* due to separate \`mmproj\` vision files. Use llama.cpp compatible backends. Do NOT use \*\*CUDA 13.2\*\* as you may get gibberish outputs. NVIDIA is working on a fix. {% endhint %} ### 🦥 Unsloth Studio Guide For this tutorial, we will be using \[Unsloth Studio\](/docs/new/studio.md), which is our new web UI for running and training LLMs. With Unsloth Studio, you can run models and input \*\*audio\*\*, image and text locally on \*\*Mac, Windows\*\*, and Linux and: {% columns %} {% column %} \* Search, download, \[run GGUFs\](/docs/new/studio.md#run-models-locally) and safetensor models \* \*\*Compare\*\* models \*\*side-by-side\*\* \* \[\*\*Self-healing\*\* tool calling\](/docs/new/studio.md#execute-code--heal-tool-calling) + \*\*web search\*\* \* \[\*\*Code execution\*\*\](/docs/new/studio.md#run-models-locally) (Python, Bash) \* \[Automatic inference\](/docs/new/studio.md#model-arena) parameter tuning (temp, top-p, etc.) \* \[Train LLMs\](/docs/new/studio.md#no-code-training) 2x faster with 70% less VRAM {% endcolumn %} {% column %} ![](https://unsloth.ai/files/dQy5izI8WRumFBHqVXtW) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Install Unsloth \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% endstep %} {% step %} #### Setup Unsloth Studio (one time) Setup automatically installs Node.js (via nvm), builds the frontend, installs all Python dependencies, and builds llama.cpp with CUDA support. {% hint style="info" %} \*\*WSL users:\*\* you will be prompted for your \`sudo\` password to install build dependencies (\`cmake\`, \`git\`, \`libcurl4-openssl-dev\`). {% endhint %} {% endstep %} {% step %} #### Launch Unsloth \*\*MacOS, Linux, WSL:\*\* \`\`\`bash source unsloth\_studio/bin/activate unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Windows Powershell:\*\* \`\`\`bash & .\\unsloth\_studio\\Scripts\\unsloth.exe studio -H 0.0.0.0 -p 8888 \`\`\` ![](https://unsloth.ai/files/J8BaejVXrezdt6B1aeUy) \*\*Then open \`http://localhost:8888\` in your browser.\*\* {% endstep %} {% step %} #### Search and download Mistral Medium 3.5 On first launch you will need to create a password to secure your account and sign in again later. Then go to the \[Unsloth Chat\](/docs/new/studio/chat.md) tab and search for Mistral 3.5 in the search bar and download your desired model and quant. {% endstep %} {% step %} #### Run Mistral 3.5 Inference parameters should be auto-set when using Unsloth Studio, however you can still change it manually. You can also edit the context length, chat template and other settings. For more information, you can view our \[Unsloth Studio inference guide\](/docs/new/studio/chat.md). {% endstep %} {% endstepper %} ### 🦙 Llama.cpp Guide For this guide we will use Unsloth Dynamic 4-bit for Mistral Medium 3.5. See: \`unsloth/Mistral-Medium-3.5-128B-GGUF\`. For these tutorials, we will use llama.cpp for fast local inference, especially if you have a CPU or high-memory unified-memory machine. \*\*1. Build llama.cpp\*\* Obtain the latest \`llama.cpp\` on GitHub. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. For Apple Mac / Metal devices, set \`-DGGML\_CUDA=OFF\`; Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` \*\*2. Run directly from Hugging Face\*\* \`\`\`bash export LLAMA\_CACHE="unsloth/Mistral-Medium-3.5-128B-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4\_K\_XL \\ --temp 0.7 \\ --chat-template-kwargs '{"reasoning\_effort":"none"}' \`\`\` For high reasoning mode: \`\`\`bash ./llama.cpp/llama-cli \\ -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4\_K\_XL \\ --temp 0.7 \\ --chat-template-kwargs '{"reasoning\_effort":"high"}' \`\`\` \*\*3. Download the model manually\*\* After installing \`huggingface\_hub\` and \`hf\_transfer\`: \`\`\`bash pip install huggingface\_hub hf\_transfer hf download unsloth/Mistral-Medium-3.5-128B-GGUF \\ --local-dir unsloth/Mistral-Medium-3.5-128B-GGUF \\ --include "\*UD-Q4\_K\_XL\*" \\ --include "\*mmproj\*" \`\`\` If downloads get stuck, set: \`\`\`bash export HF\_HUB\_ENABLE\_HF\_TRANSFER=1 \`\`\` \*\*4. Run the local GGUF\*\* \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Mistral-Medium-3.5-128B-GGUF/Mistral-Medium-3.5-128B-UD-Q4\_K\_XL.gguf \\ --temp 0.7 \\ --chat-template-kwargs '{"reasoning\_effort":"none"}' \`\`\` If a multimodal projector GGUF is included, use: \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Mistral-Medium-3.5-128B-GGUF/Mistral-Medium-3.5-128B-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/Mistral-Medium-3.5-128B-GGUF/mmproj-BF16.gguf \\ --temp 0.7 \\ --chat-template-kwargs '{"reasoning\_effort":"none"}' \`\`\` #### Llama-server deployment To deploy Mistral Medium 3.5 on llama-server, use: \`\`\`bash ./llama.cpp/llama-server \\ -hf unsloth/Mistral-Medium-3.5-128B-GGUF:UD-Q4\_K\_XL \\ --alias "mistral-medium-3.5" \\ --host 0.0.0.0 \\ --port 8001 \\ --temp 0.7 \\ --chat-template-kwargs '{"reasoning\_effort":"none"}' \`\`\` For reasoning mode: \`\`\`bash --chat-template-kwargs '{"reasoning\_effort":"high"}' \`\`\` If you're on Windows PowerShell, use: \`\`\`powershell --chat-template-kwargs "{\\"reasoning\_effort\\":\\"high\\"}" \`\`\` You can ping llama-server with an OpenAI-compatible request: \`\`\`bash curl http://localhost:8001/v1/chat/completions \\ -H "Content-Type: application/json" \\ -d '{ "model": "mistral-medium-3.5", "messages": \[ {"role": "user", "content": "Explain the main difference between instant mode and reasoning mode."} \], "temperature": 0.7 }' \`\`\` ### Mistral 3.5 Best Practices #### Prompting examples \*\*Simple reasoning prompt\*\* \`\`\` System: You are a precise reasoning assistant. Solve carefully and present only the final answer and a short explanation. User: A train leaves at 8:15 AM and arrives at 11:47 AM. How long was the journey? \`\`\` Use \`reasoning\_effort="high"\` for this style of prompt. \*\*OCR / document prompt\*\* For OCR and document extraction, put the image first and ask for structured output. \`\`\` \[image first\] Extract all text from this receipt. Return merchant, date, line\_items and total as JSON. \`\`\` \*\*Multi-modal comparison prompt\*\* \`\`\` \[image 1\] \[image 2\] Compare these two screenshots and tell me which one is more likely to confuse a new user. Give 3 concrete reasons. \`\`\` \*\*Coding agent prompt\*\* \`\`\` You are a coding agent working inside a repository. First inspect the relevant files, then propose a minimal patch. Return the final answer with: summary, files changed, tests run and risks. \`\`\` Use \`reasoning\_effort="high"\` and tool calling for codebase exploration. \*\*JSON / function calling prompt\*\* \`\`\` Use the provided tools whenever calculation or lookup is required. Return valid JSON only. Do not include prose outside the JSON object. \`\`\` ### Benchmarks ![](https://unsloth.ai/files/p3gU7s50sERw3OmC94L7) ![](https://unsloth.ai/files/sbzyGDiysPcLk4nO12CE) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/mistral-3.5.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections.md). # Connect API Providers & Model Servers to Unsloth Learn how to run models from OpenAI, Anthropic, Ollama, llama.cpp, vLLM, and other providers through a single local UI interface with \[Unsloth\](/docs/new/studio.md), an open-source repo for running and training LLMs. {% columns %} {% column %} Once connected, you can run models with code execution, tool-calling, image generation, and other features in the same Unsloth chat interface used for both local and cloud models. Unsloth uniquely supports \[prompt caching\](#prompt-caching) (to save you many tokens without accuracy degradation) while preserving access to provider-native capabilities, such as OpenAI’s built-in \[web search\](#web-search-and-thinking) and \[code execution\](#code-execution). {% endcolumn %} {% column %} {% embed url="" %} {% endcolumn %} {% endcolumns %} ### Connections Connections fall into two groups: hosted API providers that run models for you, and model servers that you run or control. \*\*Cloud Providers -\*\* Hosted APIs that use an account API key: | Connection | Capabilities | Setup guide | | ---------- | --------------------------------------- | ----------------------------------------------------------------- | | OpenAI | Image, search, code, think | \[OpenAI →\](/docs/integrations/connections/openai.md) | | Anthropic | Image, search, code, think | \[Anthropic →\](/docs/integrations/connections/anthropic-claude.md) | | OpenRouter | Many hosted models through one API key. | \[OpenRouter →\](/docs/integrations/connections/openrouter.md) | \*\*Model Servers -\*\* Inference servers running locally, on your network, or on your remote machine: | Server | Description | Guide | | --------- | ---------------------------- | --------------------------------------------------------------------------------------------------------- | | Llama.cpp | Efficient GGUF model serving | \[Llama.cpp →\](/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md) | | vLLM | High-throughput serving | \[vLLM →\](/docs/integrations/connections/vllm.md) | | Ollama | Simple local model server | \[Ollama →\](/docs/integrations/connections/ollama.md) | ### Quickstart To run an external provider's model, add an API key and select which models Unsloth should show. In this example, we’ll use \[OpenAI\](https://platform.openai.com/api-keys). The same setup works for Anthropic, and other providers. {% stepper %} {% step %} #### Create API Create a new API key from the provider’s dashboard and copy it. ![](https://unsloth.ai/files/Pmi2ri2cEhyBvFDxBNPF) {% endstep %} {% step %} #### Setup Unsloth Studio Now we will need to install and setup \[Unsloth\](/docs/new/studio.md), which will enable you to run the cloud models in a UI interface. \[See here\](/docs/new/studio/install.md) for more detailed instructions. {% tabs %} {% tab title="MacOS" %} #### Step 1: Setup Unsloth Launch the \`terminal\` from Mac, then install Unsloth by entering the command below. \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` The environment and required packages will now be installed. Type \`Y\` and press Enter when prompted to continue. After setup finishes, the server will be available locally on port \`8888\`. ![](https://unsloth.ai/files/kAxiYilqsmP233htYNpi) {% hint style="info" %} If you skipped starting the app during installation, you can launch it later with \`unsloth studio -p 8888\`. To allow connections from other devices on your network, use \`unsloth studio -H 0.0.0.0 -p 8888\` instead. {% endhint %} #### Step 2: Start Unsloth Open your browser of choice and type \`http://127.0.0.1:8888\` in the URL box. If this is your first time installing Unsloth, you will be forwarded to the Password page where you will need to create a new password. You should then see the Chat Page as shown below. ![](https://unsloth.ai/files/ryuI6lvessgKynLGfv1K) {% endtab %} {% tab title="Windows" %} #### Step 1: Setup Unsloth Open the Start Menu, search for \`PowerShell\`, and launch it. Copy & enter the install command: \`\`\`powershell irm https://unsloth.ai/install.ps1 | iex \`\`\` it will begin installing automatically. After installation finishes, PowerShell will ask if you want to start Unsloth Studio\*\*.\*\* ![](https://unsloth.ai/files/kAxiYilqsmP233htYNpi) You can also launch it with the following command: \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` {% hint style="info" %} If you would like to have your instance accessible by clients outside of your PC/computer.\\ Add \`-H 0.0.0.0\` to the \`unsloth studio\` command. {% endhint %} #### Step 2: Start Unsloth Open \`http://127.0.0.1:8888\` in your browser. On first launch, create a new password to continue to the Chat page. \*\*Unsloth Studio\*\* is now installed and ready to use. ![](https://unsloth.ai/files/ryuI6lvessgKynLGfv1K) {% endtab %} {% tab title="Linux, WSL" %} #### Step 1: Setup Unsloth {% tabs %} {% tab title="Linux" %} Open your terminal application. You can launch it by pressing \`Ctrl + Alt + T\`, or by searching for \`Terminal\` in your system's application menu. {% endtab %} {% tab title="WSL" %} Click the Windows Start Menu, type the name of your installed distro (e.g. \`Ubuntu\`), then open it. {% hint style="warning" %} On \*\*WSL\*\*, make sure your \*\*NVIDIA drivers\*\* are installed on \*\*Windows\*\* (not inside WSL) and that the \*\*CUDA toolkit\*\* is installed inside your WSL distro. See the System Requirements below for details. {% endhint %} {% endtab %} {% endtabs %} To install, copy and run the install command: \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Then: 1. Click inside the terminal window 2. Paste the command with \`Ctrl + Shift + V\` 3. Press \`Enter\` Unsloth will start setting up the environment and installing the required packages as shown below. Type \*\*Y\*\* and Press \`Enter\` when asked if you want to allow Unsloth to start now. This will start Unsloth on your local \*\*8888\*\* port. ![](https://unsloth.ai/files/uQP4sGPAd6C4MBSFdUTm) {% hint style="info" %} If you chose not to start Unsloth during the installation process, you can always start the Unsloth app using \`unsloth studio -p 8888\` . If you would like to have your Unsloth instance accessible by clients outside of your PC/computer, add \`-H 0.0.0.0\` to the \`unsloth studio\` command. {% endhint %} #### Step 2: Start Unsloth Open your browser of choice and type \`http://127.0.0.1:8888\` in the URL box. If this is your first time installing Unsloth, you will be forwarded to the Password page where you will need to create a new password. After, Unsloth should now open on the Chat Page as shown below. ![](https://unsloth.ai/files/CresfHYJ3aP1rTTlj3YF) {% endtab %} {% endtabs %} {% endstep %} {% step %} #### Configure Connections Next, connect your provider to Unsloth. 1. Open \*\*Settings\*\* → \*\*Connections\*\*, then click \*\*Add Connection.\*\* 2. Select the provider you want to add, then paste the API key you copied earlier. 3. Click \*\*Reload Models\*\* to refresh the list with models available to your account. 4. Choose the models you want to enable, then hit save. ![](https://unsloth.ai/files/GaQDk4hQbPpOhVAIBKNJ) {% endstep %} {% step %} #### Ready to Chat The models you enabled will now appear under \*\*Connected\*\* in the \*\*Select Model\*\* dropdown. ![](https://unsloth.ai/files/Q6wnsLueufwOLm7uhaEM) Unsloth dynamically exposes compatible reasoning levels and generation controls for different models. {% endstep %} {% endstepper %} ### Connect a Model Server Use this flow for \[\*\*llama.cpp\*\*\](/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server.md), \[\*\*vLLM\*\*\](/docs/integrations/connections/vllm.md), and \[\*\*Ollama\*\*\](/docs/integrations/connections/ollama.md). Start or locate the server you want to connect. {% tabs %} {% tab title="llama.cpp " %} Start \`llama-server\` with the model you want to serve: \`\`\`bash llama-server \\ --model /path/to/model.gguf \\ --host 0.0.0.0 \\ --port 8080 \`\`\` This exposes an API endpoint at: \`http://localhost:8080/v1\` To require an API key, add: \`\`\`bash --api-key 1234-myapi-key \`\`\` {% endtab %} {% tab title="vLLM" %} Start the \`vLLM\` server with the model you want to serve: \`\`\`bash vllm serve unsloth/gemma-4-26B-A4B-it \\ --dtype auto \\ \`\`\` To require an API key, add: \`\`\`bash --api-key token-abc123 \`\`\` This exposes an API endpoint at: \`http://localhost:8000/v1\` {% endtab %} {% tab title="Ollama" %} Start \`Ollama\`, then pull the model you want to use: \`\`\`bash ollama serve ollama pull qwen3:14b \`\`\` This exposes an API endpoint at: \`http://localhost:11434/v1\` {% endtab %} {% endtabs %} {% columns %} {% column %} Now we can connect the model server. Open \*\*Settings → Connections\*\*, then click \*\*Add Provider\*\*. Select llama.cpp, vLLM, or Ollama then Paste the server \*\*Base URL\*\*. \* llama.cpp example: \`http://localhost:8080/v1\` \* Ollama example: \`http://localhost:11434/v1\` {% endcolumn %} {% column %} ![](https://unsloth.ai/files/rcIydWpY9PZFkFP0tbBc) {% endcolumn %} {% endcolumns %} Click \*\*Load Models\*\* to fetch available model IDs, or enter model IDs manually if your server does not expose \`/models\`. Then, after you click \*\*Add Provider,\*\* The models you enabled will now appear under \*\*External\*\* in the \*\*Select Model\*\* dropdown. ### Code Execution When enabled, supported OpenAI and Anthropic models can run code in a provider sandbox to solve problems, analyse data, and work with files.\\ \\ Anthropic models use Claude’s provider-side Code execution tool. OpenAI uses reusable containers, which you can create, delete, and select from \*\*Code Execution\*\* settings. Select the same container in a new thread to continue with its files and state. ![](https://unsloth.ai/files/3iL48hdRoVyoL73z8DP8) \### Prompt Caching Prompt caching reduces latency and cost when requests reuse the same long prefix. It is supported for compatible providers and servers, including OpenAI, Anthropic, and llama.cpp. Use the \*\*Prompt caching\*\* setting in the side panel to control caching behaviour for supported connections. ![](https://unsloth.ai/files/rd7uqDkUz6YnddW01aRl) For llama.cpp, prompt caching is enabled by default and can be disabled when starting \`llama-server\` with: \`\`\`bash --no-cache-prompt \`\`\` ### Web Search & Thinking Provider-side web search is available for supported models from OpenAI, Anthropic, OpenRouter, Mistral, Gemini, and Kimi. The Think control adapts to the selected model: some models use an on/off toggle, while reasoning-effort models use model specific thinking levels. ![](https://unsloth.ai/files/46johOqXvwqWOhGnbso4) \### Image Generation Just like GPT and Gemini, Unsloth also supports image generation. You can directly edit an image by clicking the “Edit Image” button and entering a new prompt to refine or regenerate it. Images are generated automatically when requested, but you can toggle this behavior off. A download button is also available, allowing you to save the image in its original full resolution. ![](https://unsloth.ai/files/tc7WCuUGdy9DA5PQwvDs) ![](https://unsloth.ai/files/GgH0XVMZeBOPsDM3rQVv) \### Troubleshooting If a provider fails to connect, check that the API key belongs to the selected provider and has access to the model you chose. If a model does not appear after clicking \*\*Reload Models\*\*, it may not be available for your account. You can still use Unsloth’s default model list or choose another model. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/unsloth-start.md). # Run Coding Agents with Local LLMs using Unsloth Start Unsloth lets you connect \[Claude Code\](/docs/basics/claude-code.md), \[Codex\](/docs/basics/codex.md), Hermes, OpenCode, Pi, and other coding agents to a local model via the \`unsloth start\` command. The entire workflow can run offline on your own hardware. \[Unsloth Studio\](/docs/new/studio.md) automatically configures the endpoint, API key, provider, model, and context length for each launch, so you can use your preferred agent without modifying them. This guide will show you how to launch models from the command line for offline use, and connect to Unsloth Studio. {% columns %} {% column width="58.333333333333336%" %} ### Quickstart First, make sure you have \[Unsloth installed\](/docs/new/studio/install.md). Then open Unsloth, load a model, go to your project folder, and run the command in terminal: \`\`\`bash unsloth start claude \`\`\` You can replace \`claude\` with any agent below: {% endcolumn %} {% column width="41.666666666666664%" %} ![](https://unsloth.ai/files/HrMdgp7HDGelphJEpwCG) Claude Code running with Qwen3.5 locally. {% endcolumn %} {% endcolumns %} | Agent | Command | | ------------------------------------------------------------------ | ------------------------ | | _:claude:_ Claude Code | \`unsloth start claude\` | | _:openai:_ OpenAI Codex | \`unsloth start codex\` | | _:caduceus:_ Hermes Agent | \`unsloth start hermes\` | | _:lobster:_ OpenClaw | \`unsloth start openclaw\` | | _:rectangle-vertical:_ OpenCode | \`unsloth start opencode\` | | _:pi:_ Pi Coding Agent | \`unsloth start pi\` | Unsloth uses temporary or session-scoped provider configuration. It does not add an Unsloth provider to the agent's normal configuration files. {% hint style="info" %} Codex currently requires a GGUF model served through the \`llama-server\` backend. {% endhint %} ### Agent Guides | Title | Cover image | | | --- | --- | --- | | Claude Code | [/files/SQea9SE3lYsQ81yTgdHa](https://unsloth.ai/files/SQea9SE3lYsQ81yTgdHa) | [/pages/w020xJgdCTBtTvfHtvye](https://unsloth.ai/pages/w020xJgdCTBtTvfHtvye) | | Codex | [/files/lc9muACUvI9NGz5e4i31](https://unsloth.ai/files/lc9muACUvI9NGz5e4i31) | [/pages/PCjZ57h5pE0QccKyJMYD](https://unsloth.ai/pages/PCjZ57h5pE0QccKyJMYD) | | Hermes Agent | [/files/JWfTGA1IyyNYAoRZgDH8](https://unsloth.ai/files/JWfTGA1IyyNYAoRZgDH8) | [/pages/q1ZbCTKGY7P8eXLeDdEN](https://unsloth.ai/pages/q1ZbCTKGY7P8eXLeDdEN) | | OpenClaw | [/files/DB1O5gm73B7wigNpR3hn](https://unsloth.ai/files/DB1O5gm73B7wigNpR3hn) | [/pages/CwQEpEmkKPmyEYdnEngt](https://unsloth.ai/pages/CwQEpEmkKPmyEYdnEngt) | | OpenCode | [/files/1CeptjdIcQaih70dC3iD](https://unsloth.ai/files/1CeptjdIcQaih70dC3iD) | [/pages/qaA8ZjTxsH2GTuBOHyra](https://unsloth.ai/pages/qaA8ZjTxsH2GTuBOHyra) | | Unsloth API | [/files/kw2LlDbFk91VBhcoAc2g](https://unsloth.ai/files/kw2LlDbFk91VBhcoAc2g) | [/pages/7sCtc6YWnJBYthTjQsr7](https://unsloth.ai/pages/7sCtc6YWnJBYthTjQsr7) | \### Load a model from the command line ![unsloth start launching Codex with a local GGUF model in Unsloth Studio](https://unsloth.ai/files/WX0G6n7O1CvsZ6eVWVHL) Unsloth start finds or loads the model, configures the coding agent and launches it from the current project. You can select and load a model while launching the agent: {% tabs %} {% tab title="With quant suffix" %} \`\`\`bash unsloth start codex \\ --model unsloth/gemma-4-E2B-it-GGUF:UD-Q4\_K\_XL \\ --context-length 32768 \`\`\` The \`:UD-Q4\_K\_XL\` suffix selects the GGUF quant. {% endtab %} {% tab title="With \`--gguf-variant\`" %} \`\`\`bash unsloth start codex \\ --model unsloth/gemma-4-E2B-it-GGUF \\ --gguf-variant UD-Q4\_K\_XL \\ --context-length 32768 \`\`\` An explicit \`--gguf-variant\` overrides the quant written after the model name. {% endtab %} {% endtabs %} On the default local address, passing \`--model\` lets \`unsloth start\` start a temporary server when Unsloth Studio is not already running. The temporary server stops when the agent exits. If Unsloth is already running, the command connects to it and leaves it running. See below for examples of agents being connected to a local LLM: {% columns %} {% column width="50%" %} ![](https://unsloth.ai/files/ECtwB96Ob4AiLBLpQ3Sq) OpenCode {% endcolumn %} {% column width="50%" %} ![](https://unsloth.ai/files/gjh3Yo18OEOPBME5ip3M) Hermes {% endcolumn %} {% endcolumns %} {% columns %} {% column width="33.33333333333333%" %} ![](https://unsloth.ai/files/aDIEWsDR444JLL2lyzV9) Claude Code {% endcolumn %} {% column width="33.33333333333333%" %} ![](https://unsloth.ai/files/X5tE4Q16valzHajyCMN7) Codex {% endcolumn %} {% column %} ![](https://unsloth.ai/files/2CK4jhHmnYiV6jUZZwvh) Openclaw {% endcolumn %} {% endcolumns %} ### Connect to a remote Unsloth server Set the Unsloth URL and API key before launching an agent: \`\`\`bash export UNSLOTH\_STUDIO\_URL=https://studio.example.com export UNSLOTH\_API\_KEY=sk-unsloth-... unsloth start claude \`\`\` You can also pass the key with \`--api-key\`. For a verified local Unsloth server, \`unsloth start\` creates or reuses the API key automatically. ### Options | Option | What it does | | -------------------------------------------- | --------------------------------------------------------------------- | | \`--model\`, \`-m\` | Select a model. Without it, use the first model reported by Unsloth. | | \`--api-key\` | Supply a Unsloth API key. You can also set \`UNSLOTH\_API\_KEY\`. | | \`--launch\` / \`--no-launch\` | Launch the agent or print the generated environment and command. | | \`--serve\` / \`--no-serve\` | Allow or prevent automatic local server startup. | | \`--gguf-variant\` | Select a GGUF quantization variant. | | \`--context-length\`, \`--max-seq-length\` | Set the requested context length when loading the model. | | \`--load-in-4bit\` / \`--no-load-in-4bit\` | Control 4-bit loading for non-GGUF Hugging Face models. | | \`--tensor-parallel\` / \`--no-tensor-parallel\` | Enable or disable tensor-parallel GGUF loading for multi-GPU systems. | | \`--persist\` / \`--no-persist\` | Keep the Unsloth-managed agent home where applicable. | | \`--yolo\` | Use the selected agent's non-prompting or trust mode. | | \`-h\`, \`--help\` | Show command help. | Load options are used when Unsloth needs to load or reconcile the requested model. ### Pass normal commands to the agent Arguments that are not Unsloth options are passed to the selected agent: \`\`\`bash unsloth start claude --continue unsloth start codex --persist resume --last unsloth start opencode run --continue "Continue the previous task" unsloth start pi --persist --continue \`\`\` Use the agent's own help command for its complete list of native options. ### Sessions and \`--persist\` \`--persist\` keeps managed storage. It does not resume a conversation by itself; also pass the agent's normal resume command. | Agent | Do you need \`--persist\`? | Resume example | | --------------- | --------------------------------------------------------------------- | ---------------------------------------------------------------- | | Claude Code | No. Claude uses its normal session store. | \`unsloth start claude --continue\` | | OpenAI Codex | Yes. Its managed Codex home is temporary by default. | \`unsloth start codex --persist resume --last\` | | OpenClaw | Yes. Managed config, workspace and sessions are temporary by default. | Use \`--persist\` with the same native session ID. | | OpenCode | No. OpenCode uses its normal session store. | \`unsloth start opencode run --continue "Continue"\` | | Hermes Agent | Yes. Its managed Hermes home is temporary by default. | \`unsloth start hermes --persist --continue --oneshot "Continue"\` | | Pi Coding Agent | Yes. Its managed Pi home is temporary by default. | \`unsloth start pi --persist --continue\` | For Codex, OpenClaw, Hermes and Pi, use \`--persist\` from the first launch and again when returning to the session. {% tabs %} {% tab title="Codex" %} \`\`\`bash unsloth start codex --persist unsloth start codex --persist resume --last \`\`\` {% endtab %} {% tab title="OpenClaw" %} Use a stable session ID for a named session: \`\`\`bash unsloth start openclaw --persist \\ agent --local --session-id my-session --message "Inspect this repository" unsloth start openclaw --persist \\ agent --local --session-id my-session --message "Continue" \`\`\` {% endtab %} {% tab title="Hermes Agent" %} \`\`\`bash unsloth start hermes --persist --oneshot "Inspect this repository" unsloth start hermes --persist --continue --oneshot "Continue" \`\`\` {% endtab %} {% tab title="Pi Coding Agent" %} \`\`\`bash unsloth start pi --persist --print "Inspect this repository" unsloth start pi --persist --continue --print "Continue" \`\`\` {% endtab %} {% endtabs %} #### Print the launch command without running it \`\`\`bash unsloth start claude --no-launch \`\`\` This prints the generated environment and command. The output can contain connection credentials, so do not publish it in logs or screenshots. ### Permission-bypass mode \`--yolo\` maps to the selected agent's trust or non-prompting mode. It can reduce approval prompts and allow the agent to run commands without asking first. {% hint style="warning" %} Only use it in an environment where unrestricted agent actions are acceptable. {% endhint %} #### Common issues No Unsloth server was found Start Unsloth and load a model, or pass \\\`--model\\\` so Unsloth can start a temporary local server. Codex does not connect Use a GGUF model running through \\\`llama-server\\\`. The current Codex integration does not use the Unsloth transformers backend. A session was not restored For Codex, OpenClaw, Hermes and Pi, use \\\`--persist\\\` on the first and later launches, then pass the agent's native resume option. \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/unsloth-start.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/grpo-long-context.md). # Reinforcement Learning GRPO avec un contexte 7x plus long Le plus grand défi de l'apprentissage par renforcement (RL) est de prendre en charge de longues traces de raisonnement. Nous présentons de nouveaux algorithmes de mise en lot (batching) pour permettre \\~\*\*contexte \\~7x plus long\*\* (peut dépasser 12x) Entraînement RL sans dégradation de la précision ou de la vitesse par rapport à d'autres configurations optimisées qui utilisent FA3, des kernels et des pertes en morceaux. \* Unsloth entraîne désormais gpt-oss QLoRA avec \*\*contexte 380K\*\* sur un seul GPU NVIDIA B200 de 192 Go \* \[Qwen3\](/docs/fr/modeles/tutorials/qwen3-how-to-run-and-fine-tune.md#fine-tuning-qwen3-with-unsloth)-8B GRPO atteint \*\*contexte 110K\*\* sur un H100 à 80 Go VRAM via \[vLLM\](#vllm-for-rl) et QLoRA, et \*\*65K\*\* pour \[gpt-oss\](/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune/gpt-oss-reinforcement-learning.md) avec LoRA en BF16. \* Sur 24 Go VRAM, gpt-oss atteint 20K de contexte et 32K pour \[Qwen3-VL\](/docs/fr/modeles/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md)-8B QLoRA \* Les exécutions RL Unsloth GRPO fonctionnent avec Llama, Gemma et tous les modèles prennent automatiquement en charge des contextes plus longs Nos nouveaux kernels et algorithmes de déplacement de données et de mise en lot débloquent plus de contexte en : \* Découpage dynamique \[de séquences aplaties en morceaux\](#flattened-sequence-length-chunking) pour éviter la matérialisation de tenseurs de logits massifs et \* \[Déchargement des activations de log softmax\](#offloading-activations-for-log-softmax) ce qui empêche la croissance silencieuse de la mémoire au fil du temps. {% hint style="info" %} \*\*Vous pouvez combiner toutes les fonctionnalités d'Unsloth ensemble :\*\* 1. La \[partage de poids\](/docs/fr/commencer/reinforcement-learning-rl-guide/memory-efficient-rl.md) d'Unsloth \[vLLM\](https://github.com/vllm-project/vllm) avec \[RL économe en mémoire\](/docs/fr/commencer/reinforcement-learning-rl-guide/memory-efficient-rl.md) 2. La \[et notre fonctionnalité Standby dans\](/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md) Flex Attention \[500K Context Training\](/docs/fr/blog/500k-context-length-fine-tuning.md) 3. pour gpt-oss à long contexte et notre \[FP8 RL\](/docs/fr/commencer/reinforcement-learning-rl-guide/fp8-reinforcement-learning.md) entraînement Float8 dans \[et le\](https://unsloth.ai/blog/long-context) point de contrôle asynchrone des gradients d'Unsloth {% endhint %} ### :tada:et bien plus Premiers pas \[Notebooks GRPO\](/docs/fr/commencer/unsloth-notebooks.md#grpo-reasoning-rl-notebooks) Pour commencer, vous pouvez utiliser n'importe quel {% columns %} {% column width="33.33333333333333%" %} \[\*\*gpt-oss-20b\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) GSPO {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[\*\*(ou mettre à jour Unsloth si local) :\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) Qwen3-VL-8B {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[Qwen3-8B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_8B\_FP8\_GRPO.ipynb) Vision RL {% embed url="" %} {% endcolumn %} {% endcolumns %} GPU L4 \* \*\*Adopter Unsloth pour vos tâches RL fournit un cadre robuste pour gérer efficacement des modèles à grande échelle. Pour utiliser efficacement les améliorations d'Unsloth :\*\*Recommandations matérielles \* \*\*: Utilisation d'un NVIDIA H100 ou équivalent pour une utilisation optimale de la VRAM.\*\*Conseils de configuration \`: Assurez-vous que les paramètres\` et \`gradient\_accumulation\_steps\` batch\\\_size {% hint style="success" %} sont alignés sur vos ressources informatiques pour de meilleures performances. \`\`\` Mettez Unsloth à jour vers la dernière version sur Pypi pour obtenir les mises à jour les plus récentes : \`\`\` {% endhint %} pip install --upgrade --no-cache-dir unsloth unsloth\\\_zoo \[Nos benchmarks mettent en évidence les économies de mémoire réalisées par rapport aux versions précédentes pour GPT OSS et Qwen3-8B. Les deux graphiques ci‑dessous (sans\](/docs/fr/commencer/reinforcement-learning-rl-guide/memory-efficient-rl.md)standby \`) ont été exécutés avec\` et \`batch\_size = 4\` gradient\\\_accumulation\\\_steps=2 , puisque standby, par conception, utilise toute la VRAM. ### :1234:Pour nos benchmarks, nous comparons BF16 GRPO à Hugging Face avec toutes les optimisations activées (tous les kernels de la bibliothèque kernels, Flash Attention 3, kernels de perte en morceaux, etc.) : Découpage de longueur de séquence aplatie $$ \\text{Equation 1: } \\text{Logit Memory (GB)} = \\frac{\\text{batch size} \\times\\text{context length} \\times \\text{vocab dim}}{1024^3} $$ Précédemment, Unsloth réduisait l'utilisation mémoire du RL en évitant la matérialisation complète du tenseur de logits en découpant sur la dimension du batch. Une estimation approximative de la VRAM requise pour matérialiser les logits durant la passe avant est montrée dans l'Équation (1). \`) ont été exécutés avec\`, \`En utilisant cette formulation, une configuration avec\`, et \`context\_length = 8192\` vocab\\\_dim = 128 000 \*\*requiert environ\*\* 3,3 Go de VRAM pour stocker le tenseur de logits. \[Long Context gpt-oss\](/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md) Via \*\*l'année dernière, nous avons ensuite introduit une approche de perte fusionnée pour GRPO. Cette approche garantit qu'un seul échantillon de batch est traité à la fois, réduisant significativement l'utilisation mémoire maximale. Pour la même configuration, l'utilisation de la VRAM tombe à environ\*\*0,83 Go $$ \\text{Equation 2: }\\text{Logit Memory (GB)} = \\frac{\\text{context length} \\times \\text{vocab dim}}{1024^3} $$ ![](https://unsloth.ai/files/29779c7b35037dfd4e3dbd54a0c82bb724aeedc4) , comme reflété dans l'Équation (2). ![](https://unsloth.ai/files/b7fb4749385624735f27b1af1deb53db06596e78) Figure 1 : gpt-oss BF16 GRPO LoRA (Unsloth vs. HF avec toutes les optimisations activées) Figure 2 : Qwen3-8B QLoRA GRPO LoRA (Unsloth vs. HF avec toutes les optimisations activées) \*\*Dans cette mise à jour, nous étendons la même idée en introduisant le découpage à travers la\*\* dimension de séquence \`également. Au lieu de matérialiser les logits pour l'intégralité de l'espace\` (batch\\\_size × context\\\_length) d'un seul coup, nous aplatissons ces dimensions et les traitons en plus petits morceaux en utilisant un multiplicateur configurable. Cela permet à Unsloth de prendre en charge des contextes sensiblement plus longs sans augmenter l'utilisation mémoire maximale. \`Dans la Figure 5 ci‑dessous, nous utilisons un multiplicateur de\`max(4, context\\\_length // 4096)\`) ont été exécutés avec\`, \`En utilisant cette formulation, une configuration avec\`, \`context\_length = 8192\`, bien que n'importe quel multiplicateur puisse être spécifié selon le compromis mémoire–performance désiré. Avec ce paramètre, la même configuration d'exemple ( \*\*) nécessite désormais seulement\*\* 0,207 Go de VRAM $$ \\text{Equation 3: }\\text{Logit Memory (GB)} = \\frac{\\frac{\\text{context length}}{\\text{multiplier}} \\times \\text{vocab dim}}{1024^3} $$ ![](https://unsloth.ai/files/819b2cbd3b9d2c0c13cf800279204c73d8b3f887) pour la matérialisation des logits. ![](https://unsloth.ai/files/f0b27915e282faf75746295fd177bed595824004) Figure 3 : gpt-oss-20b (H100) Unsloth nouveau vs. ancien ![](https://unsloth.ai/files/66ab00d67bb56f33d7a399ced49df06470af940d) Figure 4 : Qwen3-8B (H100) Unsloth nouveau vs. ancien ![](https://unsloth.ai/files/13b5126387fd3b43ff499b43daa7f1a98e89261b) Figure 5 : gpt-oss-20b (H100) Figure 6 : Qwen3-8B (B200) \`Cette mise à jour est reflétée dans le\` chunked\\\_hidden\\\_states\\\_selective\\\_log\\\_softmax\`compilé ci‑dessous, qui prend désormais en charge le découpage à la fois sur les dimensions batch et séquence. Pour préserver le tenseur de logits (\`\\\[batch\\\_size, context\\\_length, vocab\\\_dim\] \`), il est toujours découpé sur la dimension batch. Le découpage supplémentaire sur la séquence est contrôlé via\` unsloth\\\_logit\\\_chunk\\\_multiplier \`Dans la Figure 5 ci‑dessous, nous utilisons un multiplicateur de\`dans la configuration GRPO ; s'il n'est pas défini, il prend par défaut \`. Dans l'exemple ci‑dessous,\` input\\\_ids\\\_chunk\\\[0\] \`\`\`python correspond à la taille des mini‑lots d'états cachés dans l'optimisation 2. logprobs\_chunk = chunked\_hidden\_states\_selective\_log\_softmax( new\_hidden\_states\_chunk, lm\_head, completion\_ids, chunks=input\_ids\_chunk.shape\[0\]\*multiplier, logit\_scale\_multiply=logit\_scale\_multiply, logit\_scale\_divide=logit\_scale\_divide, logit\_softcapping=logit\_softcapping, ) \`\`\` 1. temperature=temperature, 2. Nous utilisons torch.compile avec des options de compilation personnalisées pour réduire la VRAM et augmenter la vitesse. 3. Tous les logits découpés sont convertis en float32 pour préserver la précision. ### :ghost:Nous prenons en charge le softcapping des logits, le scaling de la température et toutes les autres fonctionnalités. Découpage des états cachés \`Nous avons également observé qu'à des longueurs de contexte plus longues, les états cachés peuvent devenir un contributeur significatif à l'utilisation mémoire. Pour la démonstration, nous supposerons\`hidden\\\_states\\\_dim=4096 $$ \\text{Hidden States Memory (GB)} = \\frac{\\text{batch size} \\times\\text{context length} \\times \\text{hidden states dim}}{1024^3} $$ . L'utilisation mémoire correspondante suit une formulation similaire au cas des logits, montrée ci‑dessous. \`Avec un\` et \`batch\_size = 8\`context\\\_length = 64000 \*\*, cela aboutirait à une utilisation de VRAM d'environ\*\*2 Go \*\*. Dans cette version, nous introduisons un découpage optionnel sur la dimension batch pour le tenseur des états cachés lors du calcul des log‑probabilités. Cela ferait que l'utilisation de la VRAM soit divisée par la taille du batch ou, dans ce cas, soit\*\*0,244 Go $$ \\text{Hidden States Memory (GB)} = \\frac{\\text{context length} \\times \\text{hidden states dim}}{1024^3} $$ . Cela réduit la VRAM maximale requise pour matérialiser les états cachés, comme reflété dans l'équation mise à jour ci‑dessous : \[500K Context Training\](/docs/fr/blog/500k-context-length-fine-tuning.md) Similaire à notre perte d'entropie croisée dans notre \*\*publication, la nouvelle implémentation\*\*ajuste automatiquement le lotage des états cachés \`. Les utilisateurs peuvent également contrôler ce comportement via\`unsloth\\\_grpo\\\_mini\\\_batch \`. Les utilisateurs peuvent également contrôler ce comportement via\` . Cependant, augmenter au‑delà de la valeur optimale peut introduire une légère augmentation des performances ou un ralentissement (généralement plus rapide) par rapport à l'ancienne fonction de perte.\`Cependant, lors d'une exécution GPT-OSS (\`context\\\_length = 8192, batch\\\_size = 4, gradient\\\_accumulation\\\_steps = 2 \`), définir\` et \`unsloth\_grpo\_mini\_batch = 1\` unsloth\\\_logit\\\_chunk\\\_multiplier = 4 \*\*entraîne\*\* peu ou pas de dégradation de la vitesse tout en réduisant l'utilisation de la VRAM d'environ 5 Go ![](https://unsloth.ai/files/90726bfd70eaf2a2c8cafa85b0062aad29322c87) {% hint style="success" %} \*\*Remarque :\*\* par rapport aux anciennes versions d'Unsloth. \`Dans les Figures 3 et 4, nous utilisons la taille de batch effective maximale, qui est 8 dans cette configuration. La taille de batch effective est calculée comme\`batch\\\_size × gradient\\\_accumulation\\\_steps \`, donnant\`4 × 2 = 8 \[. Pour une explication plus approfondie de la façon dont fonctionnent les tailles de batch effectives en RL, voir notre\](/docs/fr/commencer/reinforcement-learning-rl-guide/advanced-rl-documentation.md). {% endhint %} ### :cactus:documentation RL avancée Déchargement des activations pour le log softmax \`Lors du développement de cette version, nous avons découvert que lorsque l'on carrelait (tiling) sur la dimension batch pour les états cachés, les activations n'étaient pas déchargées après le calcul fusionné des logits et des logprobs. Comme les logits sont calculés un batch à la fois en utilisant\`hidden\\\_states\\\[i\] @ lm\\\_head , la logique existante de déchargement des activations et de checkpointing des gradients, conçue pour fonctionner dans la passe forward du modèle, ne s'appliquait pas dans ce cas. \`\`\`python Pour y remédier, nous avons ajouté une logique explicite pour décharger ces activations en dehors de la passe forward du modèle, comme montré dans le pseudocode Python ci‑dessous : class Unsloth\_Offloaded\_Log\_Softmax(torch.autograd.Function): with torch.no\_grad(): def forward(...): return output output = chunked\_hidden\_states\_selective\_log\_softmax(hidden\_states, lm\_head, ...) def backward(ctx, grad\_output): hidden\_states.requires\_grad\_(True) with torch.enable\_grad(): def forward(...): hidden\_states = ctx.saved\_hidden\_states return ... \`\`\` {% hint style="success" %} \*\*Remarque :\*\* torch.autograd.backward(output, grad\\\_output) \`Cette fonctionnalité n'est efficace que lors du découpage sur la dimension batch ou lorsque\`unsloth\\\_grpo\\\_mini\\\_batch > 1 \`), définir\`. Si tous les états cachés sont matérialisés d'un seul coup pendant la passe forward (c'est‑à‑dire, {% endhint %} ### :sparkles:), la passe backward nécessite la même quantité de mémoire sur le GPU indépendamment du fait que les activations soient déchargées. Étant donné que le déchargement des activations introduit un léger ralentissement des performances sans réduire l'utilisation mémoire dans ce cas, cela n'apporte aucun bénéfice. Configuration des paramètres : \`. Les utilisateurs peuvent également contrôler ce comportement via\` et \`), il est toujours découpé sur la dimension batch. Le découpage supplémentaire sur la séquence est contrôlé via\`Si vous ne configurez pas \*\*, nous\*\* ajusterons automatiquement ces deux paramètres \`\`\`python training\_args = GRPOConfig( ... pour vous en fonction de votre VRAM disponible et en fonction de la taille de votre longueur de contexte. Ci‑dessous toutefois comment vous pouvez modifier ces variables dans votre exécution GRPO : unsloth\_grpo\_mini\_batch = 3 ... ) \`\`\` unsloth\\\_logit\\\_chunk\\\_multiplier = 2 \`. Les utilisateurs peuvent également contrôler ce comportement via\` et \`), il est toujours découpé sur la dimension batch. Le découpage supplémentaire sur la séquence est contrôlé via\` Une visualisation des optimisations et de ![](https://unsloth.ai/files/9269ac5c2db4e958e02679cd906f7de5e5c10492) peut être vue dans le schéma ci‑dessous. \`. Les utilisateurs peuvent également contrôler ce comportement via\` Les 3 matrices représentent le lot global plus large ou \`), il est toujours découpé sur la dimension batch. Le découpage supplémentaire sur la séquence est contrôlé via\` (représenté par le nombre de crochets noirs) et les lignes de chacune des matrices représentent la longueur de contexte que le ### :vhs:découpe de la longueur de séquence (représenté par le nombre de crochets rouges). \*\*vLLM pour le RL\*\*Pour les flux de travail RL, la phase d'inférence/génération est le principal goulot d'étranglement \[vLLM\](https://github.com/vllm-project/vllm). Pour y remédier, nous utilisons , qui a accéléré la génération jusqu'à 11x par rapport à la génération normale. Depuis que GRPO a été popularisé l'année dernière, vLLM est devenu un composant central de la plupart des frameworks RL, y compris Unsloth. Nous souhaitons exprimer notre gratitude à l'équipe vLLM et à tous ses contributeurs pour leur travail, car ils jouent un rôle essentiel dans l'amélioration du RL d'Unsloth ! \[Notebooks GRPO\](/docs/fr/commencer/unsloth-notebooks.md#grpo-reasoning-rl-notebooks) Pour commencer, vous pouvez utiliser n'importe quel {% columns %} {% column width="33.33333333333333%" %} \[\*\*gpt-oss-20b\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) Pour essayer le RL à plus long contexte, vous pouvez utiliser n'importe quel {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[\*\*(ou mettre à jour Unsloth si local) :\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) Qwen3-VL-8B {% embed url="" %} {% endcolumn %} {% column width="33.33333333333333%" %} \[Qwen3-8B - \*\*FP8\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_8B\_FP8\_GRPO.ipynb) Vision RL {% embed url="" %} {% endcolumn %} {% endcolumns %} \\- GSPO --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/grpo-long-context.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/nemotron-3-nano-omni.md). # NVIDIA Nemotron 3 Nano Omni - How To Run Locally NVIDIA Nemotron-3-Nano-Omni-30B-A3B is an open 30B parameter, 3B active hybrid reasoning MoE model built for multimodal agentic workloads including \*\*audio\*\*, \*\*video\*\*, text, images and docs as input, with text output. The model runs on \*\*25GB RAM\*\* for 4-bit and 36GB for 8-bit. With a \*\*256K context\*\*, Nemotron 3 Nano Omni is the \*\*strongest omni\*\* model for its size and the highest-efficiency open multimodal model. We collaborated with NVIDIA for day zero support!\\ \*\*GGUF:\*\* \[Nemotron-3-Nano-Omni-30B-A3B-Reasoning\](https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF) ### ⚙️ Usage Guide NVIDIA recommends these settings for inference: {% columns %} {% column %} \*\*Thinking mode:\*\* \* \`temperature = 0.6\` \* \`top\_p = 0.95\` {% endcolumn %} {% column %} \*\*Instruct mode:\*\* \* \`temperature = 0.2\` {% endcolumn %} {% endcolumns %} ### Run Nemotron-3-Nano-Omni Depending on your use-case you will need to use \[different settings\](#usage-guide). Some GGUFs end up similar in size because the model architecture (like \[gpt-oss\](/docs/models/gpt-oss-how-to-run-and-fine-tune.md)) has dimensions not divisible by 128, so parts can’t be quantized to lower bits. \*\*GGUF:\*\* \[Nemotron-3-Nano-Omni-30B-A3B-Reasoning\](https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF) The 4-bit versions of the model requires \\~25GB RAM. 8-bit requires 36GB. For these guides, we will be using \`UD-Q4-K-XL\` which is a good balance between size and accuracy. [Run in Unsloth Studio](https://unsloth.ai/pages/GEABCCTb5KV7QKOeL4YY#unsloth-studio-guide) [Run in llama.cpp](https://unsloth.ai/pages/GEABCCTb5KV7QKOeL4YY#llama.cpp-tutorial) {% hint style="warning" %} Currently no multimodal/vision GGUF works in \*\*Ollama\*\* due to separate \`mmproj\` vision files. Use llama.cpp compatible backends. Do NOT use \*\*CUDA 13.2\*\* as you may get gibberish outputs. NVIDIA is working on a fix. {% endhint %} ### 🦥 Unsloth Studio Guide For this tutorial, we will be using \[Unsloth Studio\](/docs/new/studio.md), which is our new web UI for running and training LLMs. With Unsloth Studio, you can run models and input \*\*audio\*\*, image and text locally on \*\*Mac, Windows\*\*, and Linux and: {% columns %} {% column %} \* Search, download, \[run GGUFs\](/docs/new/studio.md#run-models-locally) and safetensor models \* \*\*Compare\*\* models \*\*side-by-side\*\* \* \[\*\*Self-healing\*\* tool calling\](/docs/new/studio.md#execute-code--heal-tool-calling) + \*\*web search\*\* \* \[\*\*Code execution\*\*\](/docs/new/studio.md#run-models-locally) (Python, Bash) \* \[Automatic inference\](/docs/new/studio.md#model-arena) parameter tuning (temp, top-p, etc.) \* \[Train LLMs\](/docs/new/studio.md#no-code-training) 2x faster with 70% less VRAM {% endcolumn %} {% column %} ![](https://unsloth.ai/files/dQy5izI8WRumFBHqVXtW) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Install Unsloth \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% endstep %} {% step %} #### Setup Unsloth Studio (one time) Setup automatically installs Node.js (via nvm), builds the frontend, installs all Python dependencies, and builds llama.cpp with CUDA support. {% hint style="info" %} \*\*WSL users:\*\* you will be prompted for your \`sudo\` password to install build dependencies (\`cmake\`, \`git\`, \`libcurl4-openssl-dev\`). {% endhint %} {% endstep %} {% step %} #### Launch Unsloth \*\*MacOS, Linux, WSL:\*\* \`\`\`bash source unsloth\_studio/bin/activate unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Windows Powershell:\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` ![](https://unsloth.ai/files/J8BaejVXrezdt6B1aeUy) Then open \`http://127.0.0.1:8888\` in your browser. {% endstep %} {% step %} #### Search and download NVIDIA-Nemotron-3-Nano-30B-A3B-Omni On first launch you will need to create a password to secure your account and sign in again later. Then go to the \[Unsloth Chat\](/docs/new/studio/chat.md) tab and search for Nemotron-3-Nano-Omni in the search bar and download your desired model and quant. ![](https://unsloth.ai/files/LFzX1PKl3z1FoeSCicO6) {% endstep %} {% step %} #### Run Nemotron-3-Nano-30B-A3B-Omni Inference parameters should be auto-set when using Unsloth Studio, however you can still change it manually. You can also edit the context length, chat template and other settings. For more information, you can view our \[Unsloth Studio inference guide\](/docs/new/studio/chat.md). ![](https://unsloth.ai/files/llpGZeHCgZxXgnicewTq) {% endstep %} {% endstepper %} ### 🦙 Llama.cpp Tutorial: Instructions to run in llama.cpp (note we will be using 4-bit to fit most devices): {% stepper %} {% step %} Obtain the latest \`llama.cpp\` on \[GitHub here\](https://github.com/ggml-org/llama.cpp). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} {% endstep %} {% step %} \*\*Let's first get an image!\*\* You can also upload images as well. We shall use , which is just our mini logo showing how finetunes are made with Unsloth: {% code overflow="wrap" %} \`\`\`bash wget https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/unsloth%20made%20with%20love.png -O unsloth.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/6grTWZGCnR0olP2Q8SgB) Let's get the 2nd image at {% code overflow="wrap" %} \`\`\`bash wget https://files.worldwildlife.org/wwfcmsprod/images/Sloth\_Sitting\_iStock\_3\_12\_2014/story\_full\_width/8l7pbjmj29\_iStock\_000011145477Large\_mini\_\_1\_.jpg -O picture.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/Bo2dlyVZZXbxGr2zhpKU) {% endstep %} {% step %} Now let's download the model manually. We can do this via the code below (after installing pip install huggingface\\\_hub). If downloads get stuck, see: \[Hugging Face Hub, XET debugging\](/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) {% code overflow="wrap" %} \`\`\`bash pip install huggingface\_hub hf download unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF \\ --local-dir unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF \\ --include "\*mmproj-BF16\*" \\ --include "\*UD-Q4\_K\_XL\*" # Use "\*UD-Q2\_K\_XL\*" for Dynamic 2bit \`\`\` {% endcode %} {% endstep %} {% step %} Then run the model in conversation mode: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF/mmproj-BF16.gguf \\ --temp 0.6 \\ --top-p 0.95 \\ --min-p 0.01 \`\`\` {% endcode %} {% endstep %} {% step %} You will then see the below: ![](https://unsloth.ai/files/XtkIM9SrfTiVdiXGOY0n) {% endstep %} {% step %} Then use \`/image\` to load both images in and ask "What is this image": ![](https://unsloth.ai/files/2JZdRr7q88myzkvDxKkM) ![](https://unsloth.ai/files/3XRgs6Zp87w72Fgnll9t) {% endstep %} {% step %} And for the sloth image: ![](https://unsloth.ai/files/3iSIa054DmG9TskcvVB2) {% endstep %} {% endstepper %} #### Llama-server serving & deployment To deploy Nemotron 3 Nano Omni locally, use \`llama-server\`. In a new terminal, for example via \`tmux\`, deploy the model: \`\`\`bash ./llama.cpp/llama-server \\ -hf unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:UD-Q4\_K\_XL \\ --alias "unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning" \\ --prio 3 \\ --temp 0.6 \\ --top-p 0.95 \\ --port 8001 \`\`\` If you downloaded the model manually, use: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\ --model unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF/mmproj-BF16.gguf \\ --alias "unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning" \\ --prio 3 \\ --temp 0.6 \\ --top-p 0.95 \\ --port 8001 \`\`\` {% endcode %} Then in a new terminal, after installing the OpenAI client with \`pip install openai\`: \`\`\`python from openai import OpenAI openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning", messages = \[ {"role": "user", "content": "What is 2+2?"}, \], ) print(completion.choices\[0\].message.reasoning\_content) print(completion.choices\[0\].message.content) \`\`\` Which will show something like the below: ![](https://unsloth.ai/files/zIvBGWhoBrNq0Bikh7Xh) \#### Image input through the OpenAI-compatible server Let's use \`picture.png\` which was the sloth image like in \[#llama.cpp-tutorial\](#llama.cpp-tutorial "mention") {% code expandable="true" %} \`\`\`python from openai import OpenAI import base64 import mimetypes image\_link = "picture.png" def file\_to\_data\_url(path: str) -> str: mime = mimetypes.guess\_type(path)\[0\] or "application/octet-stream" with open(path, "rb") as f: data = base64.b64encode(f.read()).decode("utf-8") return f"data:{mime};base64,{data}" openai\_client = OpenAI( base\_url = "http://127.0.0.1:8001/v1", api\_key = "sk-no-key-required", ) completion = openai\_client.chat.completions.create( model = "unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning", messages = \[ { "role": "user", "content": \[ { "type": "text", "text": "What is this image?", }, { "type": "image\_url", "image\_url": { "url": file\_to\_data\_url(image\_link), }, }, \], } \], ) print(completion.choices\[0\].message.reasoning\_content) print(completion.choices\[0\].message.content) \`\`\` {% endcode %} Which will show something like below: ![](https://unsloth.ai/files/EmvDtT1spnlaXF6A4bxz) \### 🦥 Fine-tuning Nemotron 3 Nano Omni Unsloth supports the entire \[Nemotron\](/docs/models/nemotron-3.md) model family. Nemotron 3 Nano Omni is useful for multimodal agent datasets. You can train on audio, vision or text via Unsloth. \*\*Video input\*\* fine-tuning is currently not supported. For text-only and notebooks, you can start from the existing \[Nemotron 3 Nano fine-tuning flow\](/docs/models/nemotron-3.md#fine-tuning-nemotron-3-and-rl). For multimodal adapters, make sure your dataset includes the modality your agent actually needs: \* \*\*Computer use:\*\* screenshots, UI state, cursor/context, expected next action \* \*\*Document intelligence:\*\* PDFs, screenshots, charts, tables, structured extraction targets \* \*\*Audio understanding:\*\* audio clips, sampled frames, summaries, timestamps, events and follow-up questions \* \*\*Agent loops:\*\* observation → reasoning → action → validation examples For Omni, do not blindly reuse text-only VRAM numbers. Multimodal encoders, projector weights, image tokens, audio chunks and long context all increase memory use. Start with shorter contexts and smaller batch sizes, then scale up. ### Benchmarks Nemotron 3 Nano Omni is the strongest omni model for its size. It is also the highest-efficiency open multimodal model with leading accuracy. The model surpasses Qwen3-Omni-30B-A3B on every benchmark. ![](https://unsloth.ai/files/OFcStC3QG0ZoN36kpcoA) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/nemotron-3-nano-omni.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/new/studio/chat.md). # How to Run models with Unsloth Studio \[Unsloth Studio\](/docs/new/studio.md) lets you run AI models 100% offline on your computer. Run model formats like GGUF and safetensors from Hugging Face or from your local files. \* \*\*Works on all MacOS, CPU, Windows, Linux, WSL setups! No GPU required\*\* \* \[\*\*Self-healing tool calling\*\*\](#auto-healing-tool-calling)\*\*,\*\* advanced \[\*\*web search\*\*\](#advanced-web-search), \[\*\*code execution\*\*\](#code-execution) \* Use Unsloth as an OpenAI-compatible inference \[\*\*API endpoint\*\*\](/docs/basics/api.md) or connect a \[provider\](/docs/integrations/connections.md) \* Search + Download + Run + \[Compare\](#model-arena) any model like GGUFs, LoRA adapters, safetensors etc. \* \[\*\*Auto inference parameter\*\*\](#auto-parameter-tuning) tuning (temp, top-p etc.) and edit chat templates \* Upload images, audio, PDFs, code, DOCX and more file types to chat with. ![](https://unsloth.ai/files/a6CWFBldAN3JSbnTonuH) \### Using Unsloth Studio Chat {% hint style="success" %} Unsloth Studio Chat automatically works on \*\*multi-GPU setups\*\* for inference. {% endhint %} {% columns %} {% column %} #### Code execution Unsloth Studio lets LLMs run Bash and Python, not just JavaScript. It also sandboxes programs like Claude Artifacts so models can test code, generate files, and verify answers with real computation. This makes answers from models more reliable and accurate. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/7mvEmWEU6T1SaKgv2G2A) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Auto-healing tool calling Unsloth Studio not only allows \[tool calling\](#id-50-tool-calling-accuracy), but also auto-fixes malformed or broken tool-calls by 50%. This means you'll always get inference outputs \*\*without\*\* broken tool calling. E.g. Qwen3.5-4B searched 20+ websites and cited sources, with web search happening inside its thinking trace. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/llpGZeHCgZxXgnicewTq) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Advanced Web Search Unsloth's web search actually visits pages directly to collect relevant information and data and doesn't just scan through website summaries. This provides outputs much more accurate / in-depth info and context. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/a6CWFBldAN3JSbnTonuH) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Use Unsloth as an API endpoint You can now use local LLMs via tools like \[Claude Code\](/docs/basics/claude-code.md) and \[Codex\](/docs/basics/codex.md) by connecting it to Unsloth's \[API endpoint\](#use-unsloth-as-an-api-endpoint). This means you'll be able to directly run Qwen and Gemma models in those tools with Unsloth's inference which includes features like self-healing tool-calling, websearch etc. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/Z3eIk2YCloY1lJy73JHS) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Automatic inference settings Inference parameters like \*\*temperature\*\*, \*\*top-p\*\*, \*\*top-k\*\*, \[\*\*MTP\*\*\](/docs/models/qwen3.6.md#mtp-guide) are automatically pre-set for new models like Qwen3.5 so you can get the best outputs without worrying about settings. You can also adjust parameters manually and edit the system prompt. Context length adjustment is no longer necessary with llama.cpp’s smart auto context, which uses only the context you need without loading anything extra. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/6qAnXnMmlrt4EPpQcTmC) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Connect Providers \[Unsloth connects\](/docs/integrations/connections.md) to OpenAI, Anthropic, Ollama, llama.cpp, vLLM, and others. Add API keys or model server URLs, then use external models in the same chat interface as local + cloud models. Run with \[prompt caching\](/docs/integrations/connections.md#prompt-caching), tool-calling, thinking, and provider-native features like OpenAI's \[web search\](#web-search-and-thinking) and \[code execution\](#code-execution). {% endcolumn %} {% column %} ![](https://unsloth.ai/files/6qAnXnMmlrt4EPpQcTmC) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Search and run models You can search and download any model via Hugging Face or use local files. Unsloth supports a wide range of model types, including \*\*GGUF\*\*, vision-language, and text-to-speech models. Run the latest models like \[Qwen3.5\](/docs/models/qwen3.5.md) or NVIDIA \[Nemotron 3\](/docs/models/nemotron-3.md). Upload images, audio, PDFs, code, DOCX and more file types to chat with. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/M316B4I6CcfYw32iXk3d) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} #### Chat Workspace Enter prompts, attach any documents, images (webp, png), code files, txt, or audio as additional context, and see the model’s responses in real time. Toggle on or off: Thinking + Web search. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/8cNPz5zNEK3DJljRXZkK) {% endcolumn %} {% endcolumns %} ### \*\*+50% Tool Calling Accuracy\*\* Unsloth offers several unique features that improve tool calling, including: \* Tool calls across all models in Unsloth are \*\*30% to 80% more accurate\*\*. \* Web search retrieves actual web content instead of only summaries. \* The maximum number of allowed tool calls is \*\*more than 25.\*\* \* Tool calls terminate more reliably, reducing loops and repeated calls. \* Improved tool-call healing and deduplication logic helps prevent XML from leaking into outputs. See test results with \`unsloth/Qwen3.5-4B-GGUF (UD-Q4\_K\_XL)\` with web search, code execution, and thinking enabled: | Metric | Normal Tool-calling | Unsloth Tool-calling | | ---------------------------- | ------------------- | -------------------- | | XML leaks in response | 10/10 | 0/10 | | URL fetches used | 0 | 4/10 runs | | Runs with correct song names | 0/10 | 2/10 | | Avg tool calls | 5.5 | 3.8 | | Avg response time | 12.3s | 9.8s | ### Model Arena Unsloth Chat lets you compare any two models side-by-side using the same prompt. E.g. compare the base model and LoRa adapter. Inference will firstly load for one model, then the second one (parallel inference is being worked on). ![](https://unsloth.ai/files/rPiD6ZlZKj0Kcm1xDpsD) {% columns %} {% column %} After training, you can compare the base and fine-tuned models side by side with the same prompt to see what changed and whether results improved. This workflow makes it easy to see how your fine-tuning changed the model’s responses and whether it improved results for your use case. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/jccFmdtiZyZrlAbPMyTn) {% endcolumn %} {% endcolumns %} {% hint style="success" %} Unsloth Studio Chat auto works on \*\*multi-GPU setups\*\* for inference. {% endhint %} ### Using old / existing GGUF models {% columns %} {% column %} \*\*Apr 1 update:\*\* You can now select an existing folder for Unsloth to detect from. \*\*Mar 27 update:\*\* Unsloth Studio now \*\*automatically detects older / pre-existing models\*\* downloaded from Hugging Face, LM Studio etc. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/RhH0rXxH932OE87ubTvN) {% endcolumn %} {% endcolumns %} \*\*Manual instructions:\*\* Unsloth Studio detects models downloaded to your Hugging Face Hub cache \`(C:\\Users{your\_username}.cache\\huggingface\\hub)\`. If you have GGUF models downloaded through LM Studio, note that these are stored in \`C:\\Users\\{your\_username}.cache\\lm-studio\\models\` \*\*\*OR\*\*\* \`C:\\Users{your\_username}\\lm-studio\\models\` and are not visible to llama.cpp by default - you will need to move or copy those .gguf files into your Hugging Face Hub cache directory (or another path accessible to llama.cpp) for Unsloth Studio to load them. After fine-tuning a model or adapter in Unsloth, you can export it to GGUF and run local inference with \*\*llama.cpp\*\* directly in Unsloth Chat. Unsloth Studio is powered by llama.cpp and Hugging Face. ### Adding Files as Context Unsloth Chat supports multimodal inputs directly in the conversation. You can attach documents, images, or audio as additional context for a prompt. ![](https://unsloth.ai/files/WMkZpt2taS3qU9RQmzZG) This makes it easy to test how a model handles real-world inputs such as PDFs, screenshots, or reference material. Files are processed locally and included as context for the model. ### \*\*Deleting model files\*\* You can delete old model files either from the bin icon in model search or by removing the relevant cached model folder from the default Hugging Face cache directory. By default, Hugging Face uses \`~/.cache/huggingface/hub/\` on macOS/Linux/WSL and \`C:\\Users\\\\.cache\\huggingface\\hub\\\` on Windows. \* \*\*MacOS, Linux, WSL:\*\* \`~/.cache/huggingface/hub/\` \* \*\*Windows:\*\* \`%USERPROFILE%\\.cache\\huggingface\\hub\\\` If \`HF\_HUB\_CACHE\` or \`HF\_HOME\` is set, use that location instead. On Linux and WSL, \`XDG\_CACHE\_HOME\` can also change the default cache root. ### \*\*Unsloth not detecting or using my GPU\*\* If the model is not using your GPU specifically for Docker, try: Pulling the latest image manually: \`\`\`bash docker pull unsloth/unsloth:latest \`\`\` \* Start the container with GPU access: \* \`docker run\`: \`--gpus all\` \* Docker Compose: \`capabilities: \[gpu\]\` \* On Linux, make sure the NVIDIA Container Toolkit is installed. \* On Windows: \* Check that \`nvcc --version\` matches the CUDA version shown in \`nvidia-smi\` \* Follow: \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/new/studio/chat.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide.md). # Fine-tuning LLMs Guide ## 1. What Is Fine-tuning? Fine-tuning / training / post-training models customizes its behavior, enhances + injects knowledge, and optimizes performance for domains and specific tasks. For example: \* OpenAI’s \*\*GPT-5\*\* was post-trained to improve instruction following and helpful chat behavior. \* The standard method of post training is called Supervised Fine-Tuning (SFT). Other methods include preference optimization (DPO, ORPO), distillation and \[Reinforcement Learning (RL)\](/docs/get-started/reinforcement-learning-rl-guide.md) (GRPO, GSPO), where an "agent" learns to make decisions by interacting with an environment and receiving \*\*feedback\*\* in the form of \*\*rewards\*\* or \*\*penalties\*\*. With \[Unsloth\](https://github.com/unslothai/unsloth), you can fine-tune or do RL for free on Colab, Kaggle, or locally with just 3GB VRAM by using our \[notebooks\](https://docs.unsloth.ai/get-started/unsloth-notebooks). By fine-tuning a pre-trained model on a dataset, you can: \* \*\*Update + Learn New Knowledge\*\*: Inject and learn new domain-specific information. \* \*\*Customize Behavior\*\*: Adjust the model’s tone, personality, or response style. \* \*\*Optimize for Tasks\*\*: Improve accuracy and relevance for specific use cases. \*\*Example fine-tuning or RL use-cases\*\*: \* Enables LLMs to predict if a headline impacts a company positively or negatively. \* Can use historical customer interactions for more accurate and custom responses. \* Fine-tune LLM on legal texts for contract analysis, case law research, and compliance. You can think of a fine-tuned model as a specialized agent designed to do specific tasks more effectively and efficiently. \*\*Fine-tuning can replicate all of RAG's capabilities\*\*, but not vice versa. {% columns %} {% column %} #### :question:What is LoRA/QLoRA? In LLMs, we have model weights. Llama 70B has 70 billion numbers. Instead of changing all 70B numbers, we instead add thin matrices A and B to each weight, and optimize those. This means we only optimize 1% of weights. LoRA is when the original model is 16-bit unquantized while QLoRA quantizes to 4-bit to save 75% memory. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/quSOO8Z4tamYR1YQJytM) Instead of optimizing Model Weights (yellow), we optimize 2 thin matrices A and B. {% endcolumn %} {% endcolumns %} #### Fine-tuning misconceptions: You may have heard that fine-tuning does not make a model learn new knowledge or RAG performs better than fine-tuning. That is \*\*false\*\*. You can train a specialized coding model with fine-tuning and RL while RAG can’t change the model’s weights and only augments what the model sees at inference time. Read more FAQ + misconceptions \[here\](https://unsloth.ai/docs/get-started/pages/HP82bIzgldwxWk3OSzVy#fine-tuning-vs.-rag-whats-the-difference): {% content-ref url="/pages/HP82bIzgldwxWk3OSzVy" %} \[FAQ + Is Fine-tuning Right For Me?\](/docs/get-started/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me.md) {% endcontent-ref %} > \[\*\*Introducing Unsloth Studio:\*\* \](/docs/new/studio.md) Our new open-source web UI for training and running models. This means you can now fine-tune models with no-code and have observability and automatic dataset creation features. ![](https://unsloth.ai/files/hNIArh4PhiuHNcCPtva8) \## 2. Choose the Right Model + Method If you're a beginner, it is best to start with a small instruct model like Llama 3.1 (8B) and experiment from there. You'll also need to decide between normal fine-tuning, RL, QLoRA or LoRA training: \* \*\*Reinforcement Learning (RL)\*\* is used when you need a model to excel at a specific behavior (e.g., tool-calling) using an environment and reward function rather than labeled data. We have several \[notebook examples\](/docs/get-started/unsloth-notebooks.md#grpo-reasoning-rl-notebooks), but for most use-cases, standard SFT is sufficient. \* \*\*LoRA\*\* is a parameter efficient training method that typically keeps the base model’s weights frozen and trains a small set of added low-rank adapter weights (in 16-bit precision). \* \*\*QLoRA\*\* combines LoRA with 4-bit precision to handle very large models with minimal resources. \* Unsloth also supports full fine-tuning (FFT) and pretraining, which require significantly more resources, but FFT is usually unnecessary. When done correctly, LoRA can match FFT. \* Unsloth \*\*all types models\*\*: \[text-to-speech\](/docs/basics/text-to-speech-tts-fine-tuning.md), \[embedding\](/docs/basics/embedding-finetuning.md), GRPO, RL, \[vision\](/docs/basics/vision-fine-tuning.md), multimodal and more. {% hint style="info" %} Research shows that \*\*training and serving in the same precision\*\* helps preserve accuracy. This means if you want to serve in 4-bit, train in 4-bit and vice versa. {% endhint %} We recommend starting with QLoRA, as it is one of the most accessible and effective methods for training models. Our \[dynamic 4-bit\](https://unsloth.ai/blog/dynamic-4bit) quants, the accuracy loss for QLoRA compared to LoRA is now largely recovered. ![](https://unsloth.ai/files/8Kn9Z2WF4hRnq7NN6Ex5) You can change the model name to whichever model you like by matching it with model's name on Hugging Face e.g. '\`unsloth/llama-3.1-8b-unsloth-bnb-4bit\`'. We recommend starting with \*\*Instruct models\*\*, as they allow direct fine-tuning using conversational chat templates (ChatML, ShareGPT etc.) and require less data compared to \*\*Base models\*\* (which uses Alpaca, Vicuna etc). Learn more about the differences between \[instruct and base models here\](/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use.md#instruct-or-base-model). \* Model names ending in \*\*\`unsloth-bnb-4bit\`\*\* indicate they are \[\*\*Unsloth dynamic 4-bit\*\*\](https://unsloth.ai/blog/dynamic-4bit) \*\*quants\*\*. These models consume slightly more VRAM than standard BitsAndBytes 4-bit models but offer significantly higher accuracy. \* If a model name ends with just \*\*\`bnb-4bit\`\*\*, without "unsloth", it refers to a standard BitsAndBytes 4-bit quantization. \* Models with \*\*no suffix\*\* are in their original \*\*16-bit or 8-bit formats\*\*. While they are the original models from the official model creators, we sometimes include important fixes - such as chat template or tokenizer fixes. So it's recommended to use our versions when available. There are other settings which you can toggle: \* \*\*\`max\_seq\_length = 2048\`\*\* – Controls context length. While Llama-3 supports 8192, we recommend 2048 for testing. Unsloth enables 4× longer context fine-tuning. \* \*\*\`dtype = None\`\*\* – Defaults to None; use \`torch.float16\` or \`torch.bfloat16\` for newer GPUs. \* \*\*\`load\_in\_4bit = True\`\*\* – Enables 4-bit quantization, reducing memory use 4× for fine-tuning. Disabling it enables LoRA 16-bit fine-tuning. You can also enable 16-bit LoRA with \`load\_in\_16bit = True\` \* To enable full fine-tuning (FFT), set \`full\_finetuning = True\`. For 8-bit fine-tuning, set \`load\_in\_8bit = True\`. \* \*\*Note:\*\* Only one training method can be set to \`True\` at a time. {% hint style="info" %} A common mistake is jumping straight into full fine-tuning (FFT), which is compute-heavy. Start by testing with LoRA or QLoRA first, if it won’t work there, it almost certainly won’t work with FFT. And if LoRA fails, don’t assume FFT will magically fix it. {% endhint %} You can also do \[Text-to-speech (TTS)\](/docs/basics/text-to-speech-tts-fine-tuning.md), \[reasoning (GRPO)\](/docs/get-started/reinforcement-learning-rl-guide.md), \[vision\](/docs/basics/vision-fine-tuning.md), \[RL\](/docs/get-started/reinforcement-learning-rl-guide/preference-dpo-orpo-and-kto.md) (GRPO, DPO), \[continued pretraining\](/docs/basics/continued-pretraining.md), text completion and other training methodologies with Unsloth. {% columns %} {% column %} Read our guide on choosing models: {% content-ref url="/pages/BSShKhLoFNlGWO5cN8VJ" %} \[What Model Should I Use?\](/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use.md) {% endcontent-ref %} {% endcolumn %} {% column %} For inidivudal tutorials on models: {% content-ref url="/pages/BAeSP6aOxvSeDUzCgKOK" %} \[Complete LLM Directory\](/docs/models/tutorials.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} ## 3. Your Dataset For LLMs, datasets are collections of data that can be used to train our models. In order to be useful for training, text data needs to be in a format that can be tokenized. \* You will need to create a dataset usually with 2 columns - question and answer. The quality and amount will largely reflect the end result of your fine-tune so it's imperative to get this part right. \* You can \[synthetically generate data\](/docs/get-started/fine-tuning-llms-guide/datasets-guide.md#synthetic-data-generation) and structure your dataset (into QA pairs) using ChatGPT or local LLMs. \* You can also use our new Synthetic Dataset notebook which automatically parses documents (PDFs, videos etc.), generates QA pairs and auto cleans data using local models like Llama 3.2. \[Access the notebook here.\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Meta\_Synthetic\_Data\_Llama3\_2\_\\(3B\\).ipynb) \* Fine-tuning can learn from an existing repository of documents and continuously expand its knowledge base, but just dumping data alone won’t work as well. For optimal results, curate a well-structured dataset, ideally as question-answer pairs. This enhances learning, understanding, and response accuracy. \* But, that's not always the case, e.g. if you are fine-tuning a LLM for code, just dumping all your code data can actually enable your model to yield significant performance improvements, even without structured formatting. So it really depends on your use case. \*\*\*Read more about creating your dataset:\*\*\* {% content-ref url="/pages/XgcpRfamZHmHnRHBnVE4" %} \[Datasets Guide\](/docs/get-started/fine-tuning-llms-guide/datasets-guide.md) {% endcontent-ref %} For most of our notebook examples, we utilize the \[Alpaca dataset\](https://docs.unsloth.ai/basics/tutorial-how-to-finetune-llama-3-and-use-in-ollama#id-6.-alpaca-dataset) however other notebooks like Vision will use different datasets which may need images in the answer ouput as well. ### 4. Understand Training Hyperparameters Learn how to choose the right \[hyperparameters\](/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide.md) using best practices from research and real-world experiments - and understand how each one affects your model's performance. \*\*For a complete guide on how hyperparameters affect training, see:\*\* {% content-ref url="/pages/y6obKRSk8TwyjIrCjuGE" %} \[Hyperparameters Guide\](/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide.md) {% endcontent-ref %} ## 5. Install + Requirements You can use Unsloth via two main ways, our free notebooks or locally. ### Unsloth Notebooks We would recommend beginners to utilise our pre-made \[notebooks\](/docs/get-started/unsloth-notebooks.md) first as it's the easiest way to get started with guided steps. You can later export the notebooks to use locally. Unsloth has step-by-step notebooks for \[text-to-speech\](/docs/basics/text-to-speech-tts-fine-tuning.md), \[embedding\](/docs/basics/embedding-finetuning.md), GRPO, RL, \[vision\](/docs/basics/vision-fine-tuning.md), multimodal, different use-cases and more. ### Local Installation You can also install Unsloth locally via \[Docker\](/docs/get-started/install/docker.md) or \`pip install unsloth\` (with Linux, WSL or \[Windows\](/docs/get-started/install/windows-installation.md)). Also depending on the model you're using, you'll need enough VRAM and resources. Installing Unsloth will require a Windows or Linux device. Once you install Unsloth, you can copy and paste our notebooks and use them in your own local environment. See: {% columns %} {% column %} {% content-ref url="/pages/odJXZM9jv284RKqZ2Pna" %} \[Unsloth Requirements\](/docs/get-started/fine-tuning-for-beginners/unsloth-requirements.md) {% endcontent-ref %} {% endcolumn %} {% column %} {% content-ref url="/pages/WbSfE0ITQYsNqERZwnbZ" %} \[Installation\](/docs/get-started/install.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} ## 6. Training + Evaluation Once you have everything set, it's time to train! If something's not working, remember you can always change hyperparameters, your dataset etc. You’ll see a log of numbers during training. This is the training loss, which shows how well the model is learning from your dataset. For many cases, a loss around 0.5 to 1.0 is a good sign, but it depends on your dataset and task. If the loss is not going down, you might need to adjust your settings. If the loss goes to 0, that could mean overfitting, so it's important to check validation too. ![](https://unsloth.ai/files/nrK1MtED4IGgcr5drBKN) The training loss will appear as numbers We generally recommend keeping the default settings unless you need longer training or larger batch sizes. \* \*\*\`per\_device\_train\_batch\_size = 2\`\*\* – Increase for better GPU utilization but beware of slower training due to padding. Instead, increase \`gradient\_accumulation\_steps\` for smoother training. \* \*\*\`gradient\_accumulation\_steps = 4\`\*\* – Simulates a larger batch size without increasing memory usage. \* \*\*\`max\_steps = 60\`\*\* – Speeds up training. For full runs, replace with \`num\_train\_epochs = 1\` (1–3 epochs recommended to avoid overfitting). \* \*\*\`learning\_rate = 2e-4\`\*\* – Lower for slower but more precise fine-tuning. Try values like \`1e-4\`, \`5e-5\`, or \`2e-5\`. #### Evaluation In order to evaluate, you could do manually evaluation by just chatting with the model and see if it's to your liking. You can also enable evaluation for Unsloth, but keep in mind it can be time-consuming depending on the dataset size. To speed up evaluation you can: reduce the evaluation dataset size or set \`evaluation\_steps = 100\`. For testing, you can also take 20% of your training data and use that for testing. If you already used all of the training data, then you have to manually evaluate it. You can also use automatic eval tools but keep in mind that automated tools may not perfectly align with your evaluation criteria. ## 7. Running + Deploying the model Now let's run the model after we completed the training process! You can edit the yellow underlined part! In fact, because we created a multi turn chatbot, we can now also call the model as if it saw some conversations in the past like below: ![](https://unsloth.ai/files/gcIeBxibgTdZfpz7Lprr) ![](https://unsloth.ai/files/c6ZsdgArBAnqKvZ6w6AF) Reminder Unsloth itself provides \*\*2x faster inference\*\* natively as well, so always do not forget to call \`FastLanguageModel.for\_inference(model)\`. If you want the model to output longer responses, set \`max\_new\_tokens = 128\` to some larger number like 256 or 1024. Notice you will have to wait longer for the result as well! ### Saving + Deployment For saving and deploying your model in desired inference engines like Ollama, vLLM, Open WebUI, you will need to use the LoRA adapter on top of the base model. We have designated guides for each framework: {% content-ref url="/pages/gEugERiAw2ztDNt98JVR" %} \[Inference & Deployment\](/docs/basics/inference-and-deployment.md) {% endcontent-ref %} {% columns %} {% column %} If you’re running inference on a single device (like a laptop or Mac), use llama.cpp to convert to GGUF format to use in Ollama, llama.cpp, LM Studio etc: {% content-ref url="/pages/T7ZPf3SNAwDykZNgXptE" %} \[GGUF & llama.cpp\](/docs/basics/inference-and-deployment/saving-to-gguf.md) {% endcontent-ref %} {% endcolumn %} {% column %} If you’re deploying an LLM for enterprise or multi-user inference for FP8, AWQ, use vLLM: {% content-ref url="/pages/fhJtaLFFXVsGnbMUiACo" %} \[vLLM\](/docs/basics/inference-and-deployment/vllm-guide.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} We can now save the fine-tuned model as a small 100MB file called a LoRA adapter like below. You can instead push to the Hugging Face hub as well if you want to upload your model! Remember to get a Hugging Face \[token\](https://huggingface.co/settings/tokens) and add your token! ![](https://unsloth.ai/files/ckdSoKSGiZHxuYsmRLGX) ![](https://unsloth.ai/files/9r0QCOEc1oEapNdThVDc) After saving the model, we can again use Unsloth to run the model itself! Use \`FastLanguageModel\` again to call it for inference! ## 8. We're done! You've successfully fine-tuned a language model and exported it to your desired inference engine with Unsloth! To learn more about fine-tuning tips and tricks, head over to our blogs which provide tremendous and educational value: If you need any help on fine-tuning, you can also join our Discord server \[here\](https://discord.gg/unsloth) or \[Reddit r/unsloth\](https://www.reddit.com/r/unsloth/). Thanks for reading and hopefully this was helpful! ![](https://unsloth.ai/files/7rypIJQYTTmfdA6igk1d) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/fine-tuning-llms-guide.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/mtp.md). # How to Run MTP Models: Multi-Token Prediction Guide MTP, or Multi-Token Prediction, speeds up inference by letting a model predict multiple upcoming tokens at once instead of generating one token per step. It enables faster inference without accuracy loss and is especially effective on GPUs. In this guide, you’ll learn how to use MTP models like \[Gemma 4\](/docs/models/gemma-4.md) or \[Qwen3.6\](/docs/models/qwen3.6.md) on your local device. MTP predicts multiple future tokens, which the main model verifies in parallel. This reduces generation forward passes, speeding output while preserving quality because only verified tokens are kept. When running \[GGUFs\](/docs/basics/unsloth-dynamic-2.0-ggufs.md), MTP can make generation \*\*\\~1.4× to 2.2× faster\*\*. Dense models like Gemma-4-31B benefit most, reaching \*\*>1.4× speedup\*\* over the original. Gains are smaller on devices with lower memory bandwidth, such as older Macs. You can run MTP models directly in \[Unsloth Studio’s UI\](/docs/new/studio.md) or llama.cpp. {% hint style="info" %} \*\*MTP uses more memory than standard\*\*, so plan for \\~2 GB additional RAM/VRAM headroom. {% endhint %} [Gemma 4 MTP](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#gemma-4-mtp) [Qwen3.6 MTP](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#qwen3.6-mtp) We found \`--spec-draft-n-max 2\` is the best starting point however, \*\*do not assume \`2\` is optimal\*\*, as performance is hardware-dependent. Try any value from \`1\` through \`6\` and use whichever is fastest for your system. Unsloth Studio automatically sets the ideal MTP settings optimized for your specific hardware (Mac, CPU, GPU etc.) - you can still change it later. ### Gemma 4 MTP Google DeepMind trained MTP separately from the original \[Gemma 4\](/docs/models/gemma-4/qat.md) models, including for \[QAT variants\](/docs/models/gemma-4/qat.md). Unlike Qwen, Google released specific MTP variants under the \`assistant\` name. For best results, we only upload 3 precision options: \*\*8-bit\*\* and \*\*16-bit\*\* (BF16, F16). For QAT - we applied the \[smart 4-bit recovery process\](/docs/models/gemma-4/qat.md#qat-analysis) like we did for Gemma 4 QAT quants, and so the MTP quants are also smart 4-bit derived. We uploaded \`mtp-\` prefixed GGUFs to each repo, so you only need to use the \*\*regular original Gemma 4 GGUFs\*\*, no separate repo is needed. You can access Gemma \[MTP models here\](https://huggingface.co/collections/unsloth/gemma-4) and they can now run in \[Unsloth\](#unsloth-studio-mtp-guide). We benchmarked Gemma 4 QAT with MTP, and it runs 1.5x - 2.2x faster: ![](https://unsloth.ai/files/1dEDbzUvXjmTZeeol2VE) \*\*Table: MTP hardware requirements\*\* (units = total memory: RAM + VRAM, or unified memory) | Gemma 4 variant | 4-bit | 8-bit | BF16 / FP16 | | --------------- | -------: | -------: | ----------: | | \*\*E2B\*\* | 5 GB | 6–9 GB | 11 GB | | \*\*E4B\*\* | 6.5–7 GB | 10–13 GB | 17 GB | | \*\*12B Unified\*\* | 8–9 GB | 14–15 GB | 26 GB | | \*\*26B A4B\*\* | 17–18 GB | 29–31 GB | 53 GB | | \*\*31B\*\* | 18–21 GB | 35–39 GB | 63 GB | {% hint style="warning" %} \*\*Gemma 4 MTP is automatically enabled in\*\* \[\*\*Unsloth Studio\*\*\](#unsloth-studio-mtp-guide)\*\*. You only need to download the regular original Gemma 4 GGUFs.\*\* We updated the Gemma 4 GGUF files to include an additional MTP file inside a separate folder within the GGUF package, so there is no need to download a separate Gemma 4 assistant GGUF. The only model that still requires a separate MTP GGUF is Qwen3.6. {% endhint %} To run the Gemma 4 MTP models, follow the steps either for \[Unsloth Studio\](#unsloth-studio-mtp-guide) or \[llama.cpp\](#llama.cpp-mtp-guide). [🦥 Run in Unsloth Studio](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#unsloth-studio-mtp-guide) [🦙 Run in llama.cpp](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#llama.cpp-mtp-guide) so the below just works (this uses the 8-bit one) {% code overflow="wrap" %} \`\`\`bash llama-server \\ -hf unsloth/gemma-4-31B-it-GGUF \\ --spec-type draft-mtp \\ --spec-draft-n-max 4 \`\`\` {% endcode %} ### Qwen3.6 MTP Qwen directly trained MTP inside of the \[Qwen3.6\](/docs/models/qwen3.6.md) and \[Qwen3.5\](/docs/models/qwen3.5.md) models. This enables Qwen3.6 27B MTP to reach 160 tokens/s and Qwen3.6 35B-A3B reach 240 tokens/s on an RTX 6000 GPU. GGUF uploads: | \[Qwen3.6-27B-MTP-GGUF\](https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF) | \[Qwen3.6-35B-A3B-MTP-GGUF\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF) | | --------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | \*\*Table: MTP hardware requirements\*\* (units = total memory: RAM + VRAM, or unified memory) | Qwen3.6 | 3-bit | 4-bit | 6-bit | 8-bit | BF16 | | --- | --- | --- | --- | --- | --- | | **27B** | 16 GB | 19 GB | 25 GB | 31 GB | 56 GB | | **35B-A3B** | 18 GB | 24 GB | 31 GB | 39 GB | 71 GB | Below are graphs of inference throughput for MTP vs. no MTP: ![](https://unsloth.ai/files/PcJYNAL2D5V189UKVHV9) ![](https://unsloth.ai/files/2zkvs1iYgzwBfLxGi6Ap) We also \[uploaded MTP GGUFs\](https://huggingface.co/unsloth/models?search=mtp) for the \[Qwen3.5\](/docs/models/qwen3.5.md) \*\*model family\*\* including: 0.8B, 2B, 4B, 9B, 27B, 35B-A3B, 122B-A10B and 397B-A17B. Llama.cpp is continually improving MTP performance, so expect it to get faster overtime! To run the Qwen MTP models, follow the steps either for \[Unsloth Studio\](#unsloth-studio-mtp-guide) or \[llama.cpp\](#llama.cpp-mtp-guide). ### 🦥 Unsloth Studio MTP Guide Unsloth Studio automatically sets the ideal MTP settings optimized for your specific hardware (Mac, CPU, GPU etc.) - you can still change it later. {% stepper %} {% step %} #### Install Unsloth Run in your terminal: \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% endstep %} {% step %} #### Launch Unsloth \*\*MacOS, Linux, WSL and Windows:\*\* \`\`\`bash unsloth studio -H 127.0.0.1 -p 8888 \`\`\` Then open \`http://127.0.0.1:8888\` (or your specific URL) in your browser. {% endstep %} {% step %} #### Search and download your desired model On first launch you will need to create a password to secure your account and sign in again later. Then go to the \[Unsloth Chat\](/docs/new/studio/chat.md) tab and search for Qwen3.6 MTP or Gemma 4 in the search bar and download your desired model and quant. {% hint style="warning" %} \*\*Gemma 4 MTP is automatically enabled in Unsloth. You only need to download the regular original Gemma 4 GGUF.\*\* We updated the Gemma 4 GGUF files to include an additional MTP file inside a separate folder within the GGUF package, so there is no need to download a separate Gemma 4 assistant GGUF. The only model that still requires a separate MTP GGUF is Qwen3.6. {% endhint %} ![](https://unsloth.ai/files/mvaV201dhzJiQuSroh4E) ![](https://unsloth.ai/files/X2vsCuTdYdpQNQ6ZIMB6) {% endstep %} {% step %} #### Run your MTP model Inference, MTP and speculative \*\*decoding settings\*\* should be auto-set when using Unsloth Studio, however you can still change it manually. You can also edit speculative decoding, the context length, chat template and other settings in the right side bar. ![](https://unsloth.ai/files/lC44po1CMW2mjJYLb2G5) For more information, you can view our \[Unsloth Studio inference guide\](/docs/new/studio/chat.md). Below, the 2-bit Qwen3.6 MTP GGUF made 10+ tool calls, searched 10 sites and executed Python code: ![](https://unsloth.ai/files/GpNoIzyrR7boop0DbLNf) {% endstep %} {% endstepper %} ### 🦙 Llama.cpp MTP Guide {% stepper %} {% step %} Install the latest version of \`llama.cpp\` on \[\*\*GitHub here\*\*\](https://github.com/ggml-org/llama.cpp/pull/22673). You can follow the build instructions below as well. Change \`-DGGML\_CUDA=ON\` to \`-DGGML\_CUDA=OFF\` if you don't have a GPU or just want CPU inference. \*\*For Apple Mac / Metal devices\*\*, set \`-DGGML\_CUDA=OFF\` then continue as usual - Metal support is on by default. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} If you want to use \`llama.cpp\` directly to load models, you can do the below: (:\`Q4\_K\_XL\`) is the quantization type. You can also download via Hugging Face (point 3). This is similar to \`ollama run\` . Use \`export LLAMA\_CACHE="folder"\` to force \`llama.cpp\` to save to a specific location. The model has a maximum of 256K context length. Follow one of the commands for the specific models: [Gemma 4](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#gemma-4-mtp-1) [Qwen3.6](https://unsloth.ai/pages/3PWlU172DOGeqxIflfP7#qwen3.6-mtp-1) #### Gemma 4 MTP: Don't forget to \*\*change the model name\*\* to your desired Gemma 4 model size like Gemma-4-26B-A4B etc. as the instructions below are for Gemma-4-12B. Notice we provided a \`mtp-\` prefixed GGUF, so the below \`-hf\` command should auto download and use MTP. \*\*Thinking mode:\*\* \`\`\`bash export LLAMA\_CACHE="unsloth/gemma-4-12b-it-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4\_K\_XL \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 64 \\ --spec-type draft-mtp --spec-draft-n-max 2 \`\`\` {% hint style="info" %} Please see Gemma 4's new \[Preserved Thinking\](#thinking-enable-disable--preserve-thinking). {% endhint %} \*\*Non-thinking mode\*\*: \`\`\`bash export LLAMA\_CACHE="unsloth/gemma-4-12b-it-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4\_K\_XL \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 64 \\ --spec-type draft-mtp --spec-draft-n-max 2 \\ --chat-template-kwargs '{"enable\_thinking":false}' \`\`\` #### Qwen3.6 MTP: Don't forget to \*\*change the model name\*\* to your desired Qwen3.6 variant like Qwen3.6-35B-A3B or Qwen3.5 etc. as the instructions below are for Qwen3.6-27B: \*\*Thinking mode\*\* (General tasks)\*\*:\*\* \`\`\`bash export LLAMA\_CACHE="unsloth/Qwen3.6-27B-MTP-GGUF" ./llama.cpp/llama-cli \\ -hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q4\_K\_XL \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 20 \\ --min-p 0.00 \\ --spec-type draft-mtp --spec-draft-n-max 2 \`\`\` For precise coding tasks, change: \`temperature=0.6\` {% hint style="info" %} Please see Qwen3.6's new \[Preserved Thinking\](#thinking-enable-disable--preserve-thinking). {% endhint %} \*\*Non-thinking mode\*\* (General tasks): \`\`\`bash export LLAMA\_CACHE="unsloth/Qwen3.6-27B-MTP-GGUF" ./llama.cpp/llama-server \\ -hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q4\_K\_XL \\ --temp 0.7 \\ --top-p 0.8 \\ --top-k 20 \\ --presence-penalty 1.5 \\ --min-p 0.00 \\ --spec-type draft-mtp --spec-draft-n-max 2 \\ --chat-template-kwargs '{"enable\_thinking":false}' \`\`\` {% endstep %} {% step %} #### Manually downloading quants If you want to manually download the quants and the MTP quants, you can also do that! Download the model via the code below (after installing \`pip install huggingface\_hub hf\_transfer\`). You can choose Q4\\\_K\\\_M or other quantized versions like \`UD-Q4\_K\_XL\` . We recommend using at least 2-bit dynamic quant \`UD-Q2\_K\_XL\` to balance size and accuracy. If downloads get stuck, see: \[Hugging Face Hub, XET debugging\](/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) #### Gemma 4 MTP: \`\`\`bash hf download unsloth/gemma-4-12B-it-qat-GGUF \\ --local-dir unsloth/gemma-4-12B-it-qat-GGUF \\ --include "\*mmproj-F16\*" \\ --include "mtp-\*" \\ --include "\*UD-Q4\_K\_XL\*" # Use "\*UD-Q2\_K\_XL\*" for Dynamic 2bit \`\`\` #### Qwen3.6 MTP: \`\`\`bash hf download unsloth/Qwen3.6-27B-MTP-GGUF \\ --local-dir unsloth/Qwen3.6-27B-MTP-GGUF \\ --include "\*mmproj-F16\*" \\ --include "\*UD-Q4\_K\_XL\*" # Use "\*UD-Q2\_K\_XL\*" for Dynamic 2bit \`\`\` {% endstep %} {% step %} Then run the model in conversation mode: #### Gemma 4 MTP: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/gemma-4-12B-it-qat-GGUF/gemma-4-12B-it-qat-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/gemma-4-12B-it-qat-GGUF/mmproj-F16.gguf \\ --model-draft unsloth/gemma-4-12B-it-qat-GGUF/mtp-gemma-4-12B-it.gguf \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 64 \\ --spec-type draft-mtp --spec-draft-n-max 2 \`\`\` {% endcode %} And you will see the below - ignore the error messages as well ![](https://unsloth.ai/files/JW0VTfbCDcu5ScLOCP4x) #### Qwen3.6 MTP: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/Qwen3.6-27B-MTP-GGUF/Qwen3.6-27B-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/Qwen3.6-27B-MTP-GGUF/mmproj-F16.gguf \\ --temp 1.0 \\ --top-p 0.95 \\ --min-p 0.00 \\ --top-k 20 \\ --spec-type draft-mtp --spec-draft-n-max 2 \`\`\` {% endcode %} {% endstep %} {% step %} #### Llama-server deployment To deploy Gemma-4 on llama-server, use: {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\ --model unsloth/gemma-4-12B-it-qat-GGUF/gemma-4-12B-it-qat-UD-Q4\_K\_XL.gguf \\ --mmproj unsloth/gemma-4-12B-it-qat-GGUF/mmproj-F16.gguf \\ --model-draft unsloth/gemma-4-12B-it-qat-GGUF/mtp-gemma-4-12B-it.gguf \\ --temp 1.0 \\ --top-p 0.95 \\ --top-k 64 \\ --alias "unsloth/gemma-4-12b-it-qat-GGUF" \\ --port 8001 \\ --chat-template-kwargs '{"enable\_thinking":true}' \`\`\` {% endcode %} {% endstep %} {% endstepper %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/mtp.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/training-ai-agents-with-rl.md). # Entraîner des agents IA avec du RL L'IA « agentique » devient de plus en plus populaire au fil du temps. Dans ce contexte, un « agent » est un LLM auquel on donne un objectif de haut niveau et un ensemble d'outils pour l'atteindre. Les agents sont aussi généralement « multi‑tours » — ils peuvent effectuer une action, voir quel effet elle a eu sur l'environnement, puis effectuer une autre action de façon répétée, jusqu'à ce qu'ils atteignent leur objectif ou échouent en essayant. Malheureusement, même des LLM très performants peuvent avoir du mal à accomplir de manière fiable des tâches agentiques complexes et multi‑tours. Fait intéressant, nous avons découvert que former des agents en utilisant un algorithme RL appelé \[GRPO (Group Relative Policy Optimization)\](/docs/fr/commencer/reinforcement-learning-rl-guide/tutorial-train-your-own-reasoning-model-with-grpo.md) peut les rendre bien plus fiables ! Dans ce guide, vous apprendrez comment construire des agents IA fiables en utilisant des outils open‑source. ## 🎨 Former des agents RL avec ART \[ART (Agent Reinforcement Trainer)\](https://github.com/openpipe/art) construit au‑dessus de \[Unsloth\](https://github.com/unslothai/unsloth)Le GRPOTrainer de , est un outil qui rend la formation d'agents multi‑tours possible et facile. Si vous utilisez déjà Unsloth pour GRPO et devez former des agents capables de gérer des interactions complexes et multi‑tours, ART simplifie le processus. ![](https://unsloth.ai/files/152d6153fdb3c5e24e99418d75a95fb9d85ab4a9) Les modèles d'agents entraînés avec Unsloth+ART sont souvent capables de surpasser les modèles basés sur des prompts sur des flux de travail agentiques. \### ART + Unsloth ART s'appuie sur l'implémentation GRPO de Unsloth, efficace en mémoire et en calcul. De plus, il ajoute les fonctionnalités suivantes : #### 1. Formation d'agents multi‑tours ART introduit le concept de « trajectoire », qui se construit au fur et à mesure que votre agent s'exécute. Ces trajectoires peuvent ensuite être notées et utilisées pour GRPO. Les trajectoires peuvent être complexes et inclure même des historiques non linéaires, des appels à des sous‑agents, etc. Elles prennent également en charge les appels d'outils et les réponses. #### 2. Intégration flexible dans des bases de code existantes Si vous avez déjà un agent fonctionnant avec un modèle par prompt, ART essaie de minimiser le nombre de modifications nécessaires pour encapsuler votre boucle d'agent existante et l'utiliser pour l'entraînement. Architecturalement, ART est divisé en un client « frontend » qui vit dans votre base de code et communique via API avec un « backend » où l'entraînement réel a lieu (ceux‑ci peuvent aussi être colocalisés sur une seule machine si vous préférez utiliser le \`LocalBackend\`). Cela apporte quelques avantages clés : \* \*\*Configuration minimale requise\*\*: Le frontend ART a des dépendances minimales et peut être facilement ajouté à des bases de code Python existantes. \* \*\*S'entraîner depuis n'importe où\*\*: Vous pouvez exécuter le client ART sur votre ordinateur portable et laisser le serveur ART lancer un environnement éphémère avec GPU, ou exécuter sur un GPU local \* \*\*API compatible OpenAI\*\*: Le backend ART expose votre modèle en cours d'entraînement via une API compatible OpenAI, ce qui est compatible avec la plupart des bases de code existantes. #### 3. RULER : Récompenses d'agent en zero‑shot ART fournit également une fonction de récompense générale intégrée appelée \[RULER\](https://art.openpipe.ai/fundamentals/ruler) (Relative Universal LLM‑Elicited Rewards), qui peut éliminer le besoin de fonctions de récompense conçues à la main. De façon surprenante, les agents entraînés par RL avec la fonction de récompense automatique RULER égalent souvent ou dépassent les performances des agents entraînés avec des fonctions de récompense écrites manuellement. Cela facilite la prise en main du RL. ![](https://unsloth.ai/files/baa362bfa41cda2cd8420b1c58a888a3a7860c38) \`\`\`python # Avant : des heures d'ingénierie des récompenses def complex\_reward\_function(trajectory): # Plus de 50 lignes de logique de notation soigneuse... pass # Après : Une seule ligne avec RULER judged\_group = await ruler\_score\_group(group, "openai/o3") \`\`\` ### Quand choisir ART ART peut convenir aux projets qui nécessitent : 1. \*\*Capacités d'agent en plusieurs étapes\*\*: Lorsque votre cas d'utilisation implique des agents qui doivent effectuer plusieurs actions, utiliser des outils ou entretenir des conversations prolongées 2. \*\*Prototypage rapide sans ingénierie des récompenses\*\*: La notation automatique des récompenses par RULER peut réduire le temps de développement de votre projet de 2 à 3 fois 3. \*\*Intégration avec des systèmes existants\*\*: Lorsque vous devez ajouter des capacités RL à une base de code agentique existante avec un minimum de modifications ### Exemple de code : ART en action \`\`\`python import art from art.rewards import ruler\_score\_group # Initialiser le modèle avec un modèle de base pris en charge par Unsloth model = art.TrainableModel( name="agent-001", project="my-agentic-task", base\_model="Qwen/Qwen2.5-14B-Instruct", # Tout modèle pris en charge par Unsloth ) # Définissez votre fonction de rollout async def rollout(model: art.Model, scenario: Scenario) -> art.Trajectory: openai\_client = model.openai\_client() trajectory = art.Trajectory( messages\_and\_choices=\[ {"role": "system", "content": "..."}, {"role": "user", "content": "..."} \] ) # Votre logique d'agent ici... return trajectory # Entraînez avec RULER pour des récompenses automatiques groups = await art.gather\_trajectory\_groups( ( art.TrajectoryGroup(rollout(model, scenario) for \_ in range(8)) for scenario in scenarios ), after\_each=lambda group: ruler\_score\_group( group, "openai/o3", swallow\_exceptions=True ) ) await model.train(groups) \`\`\` ### Prise en main Pour ajouter ART à votre projet basé sur Unsloth : \`\`\`bash pip install openpipe-art # ou \`uv add openpipe-art\` \`\`\` Puis consultez les \[notebooks d'exemple\](https://art.openpipe.ai/getting-started/notebooks) pour voir ART en action avec des tâches telles que : \* Agents de récupération d'e-mails qui battent o3 \* Agents jouant à des jeux (2048, Tic Tac Toe, Codenames) \* Tâches de raisonnement complexes (Temporal Clue) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/reinforcement-learning-rl-guide/training-ai-agents-with-rl.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/api.md). # How to use Unsloth as an API endpoint You can run \*\*local LLMs\*\* with tools like \[Claude Code\](/docs/basics/claude-code.md) and \[Codex\](/docs/basics/codex.md) by connecting those tools to Unsloth’s \*\*OpenAI-compatible API endpoint\*\*. This lets you run models like \[Qwen\](/docs/models/qwen3.6.md) and \[Gemma\](/docs/models/gemma-4.md) locally for agentic coding. Unsloth also has beneficial features such as self-healing \*\*tool calling\*\*, \*\*code execution\*\*, and \*\*web search\*\*. Unsloth makes it easy to deploy a fast API inference endpoint that provides: \* \[\*\*Self-healing tool calling\*\*\](/docs/new/studio/chat.md#auto-healing-tool-calling), which helps reduce broken or malformed tool calls by 50% \* \[\*\*Code execution\*\*\](/docs/new/studio/chat.md#code-execution) support, allowing Bash and Python execution for more accurate code outputs. \* \*\*Advanced\*\* \[\*\*Web search\*\*\](/docs/new/studio/chat.md#advanced-web-search) that visits and actually reads webpages to gather in-depth info. \* \[\*\*Automatic inference\*\* settings\](/docs/new/studio/chat.md#auto-parameter-tuning) for GGUF models (temp, top-k etc.) {% columns %} {% column %} Models loaded in Unsloth (including GGUFs) are exposed as an \*\*authenticated API\*\* via \`llama-server\`. A long API key is generated for security reasons like how OpenAI provides one. Your \*\*local models\*\* can then be used directly in your preferred AI agent, SDK, or chat client. Unsloth speaks two dialects on the same port. Both support streaming, tool calling (OpenAI \`tools\` / Anthropic \`tools\`), and vision inputs: {% endcolumn %} {% column %} ![](https://unsloth.ai/files/Z3eIk2YCloY1lJy73JHS) {% endcolumn %} {% endcolumns %} \* \*\*Anthropic-compatible \`/v1/messages\`\*\* for Claude Code, OpenClaw, the Anthropic SDK, and any client that expects the Messages API. \* \*\*OpenAI-compatible \`/v1/chat/completions\`\*\* and \*\*\`/v1/responses\`\*\* for the OpenAI SDK, OpenCode, Cursor, Continue, Cline, Open WebUI, SillyTavern, and any OpenAI-compatible tool. ### ⚡ Quickstart 1. \*\*Install or update\*\* \[\*\*Unsloth Studio\*\*\](/docs/new/studio.md)\*\*.\*\* Then launch Unsloth. 2. \*\*Load a model.\*\* Click \*\*New Chat\*\*, pick or search a model (GGUF), and wait for it to finish loading. 3. \*\*Create an API key.\*\* Click your \*\*Unsloth\*\* avatar in the bottom-left → \*\*Settings\*\* → \*\*API\*\* → type a key name → \*\*Create\*\*. Copy the \`sk-unsloth-…\` value that appears. Unsloth only shows it once. 4. \*\*Point your client at Unsloth.\*\* Use \`http://localhost:PORT\` as the base URL and your \`sk-unsloth-…\` key for auth. Jump to the recipe for your tool below. ### 🔑 Creating an API key 1. Open the sidebar, click your \*\*Unsloth\*\* avatar at the bottom-left. 2. Go to \*\*Settings\*\* → \*\*API\*\* (globe :globe\\\_with\\\_meridians: icon). 3. Enter a friendly name (e.g. \`claude-code-macbook\`). Set an expiry (optional) 4. Click \*\*Create\*\*. 5. \*\*Copy the key.\*\* Unsloth stores only a hash and you won't be able to view it again. ![](https://unsloth.ai/files/mIewhCcJSWNVe9g92qw6) All keys start with the \`sk-unsloth-\` prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with \`401 Unauthorized\`. {% hint style="warning" %} Treat your API key like a password. Anyone with the key and network access to your Unsloth instance can send requests to your loaded model. {% endhint %} ### ⏳ Model Loading {% stepper %} {% step %} #### Select Model Before using the API, load a model from the \*\*Select model\*\* dropdown in the top-left corner of the Chat page. ![](https://unsloth.ai/files/qjw3MAiRKvLO7rcVlB9r) In this guide, we’ll use: \`unsloth/gemma-4-26B-A4B-it-GGUF\` with the recommended \`UD-Q4\_K\_XL\` quantization. {% endstep %} {% step %} #### Test the Model Before using the Client, send a quick message: ![](https://unsloth.ai/files/QXM0lfihazCXbxesdwDi) {% hint style="info" %} This confirms that the model loaded correctly and is ready to respond. {% endhint %} {% endstep %} {% step %} #### \*\*Unsloth API key\*\* In Unsloth, open \*\*Settings → API\*\* to view or create your API key. ![](https://unsloth.ai/files/lnHH6JFk2bFBG8Nh96hf) Treat your API key like a password and avoid exposing it in screenshots or repositories. {% endstep %} {% endstepper %} ### _:terminal:_ Unsloth run command 1. \*\*Install or update Unsloth Studio.\*\* Earlier versions don't expose the external API. See Installation. 2. \*\*Load a GGUF model.\*\* load a GGUF model using the run command. This will also load the UI on the default port. The endpoint URL and API Key will be printed out to the console , ready for you to be used with your client of choice. \`\`\`bash unsloth run --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \`\`\` #### Loading a model from the CLI You can load a model and have an API key created for you automatically using the \`unsloth\` CLI tool. When the model finishes loading, the endpoint URL and API key are printed to your console. Copy them into your client of choice and you're ready to go. #### Before you start Make sure you're on a recent version of Unsloth Studio as earlier versions don't expose the external API. See \[installation\](/docs/new/studio/install.md). #### The quick way Open a terminal and load a GGUF model: \`\`\`bash unsloth run --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \`\`\` This starts the server on the default port, loads the UI, and prints your endpoint URL and API key. #### How the model name works You can point at a model in a few different ways. Pick the one you find easiest: \`\`\`bash # Combined: repo and quantization variant in one string (recommended — shortest) unsloth run --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL # Separate: repo and variant as two flags (the older style, still works) unsloth run --model unsloth/gemma-4-26B-A4B-it-GGUF --gguf-variant UD-Q4\_K\_XL # Using -hf / --hf-repo (matches llama.cpp's spelling, handy if you're coming from there) unsloth run -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \`\`\` ### Tuning the run (optional) You don't need any of this for a basic load, but \`unsloth run\` supports many llama-server runtime flags for customizing performance, memory usage, context length, generation behavior, networking, and tool access. Additional flags are forwarded directly to the underlying inference server, and your values override Unsloth's defaults. #### Adjust generation behavior Sampling settings control how creative, focused, or deterministic the model behaves during generation. \`\`\`bash # Lower randomness and improve reproducibility unsloth run \\ --model unsloth/Qwen3-1.7B-GGUF \\ --temp 0.6 \\ --seed 42 \`\`\` Lower temperature values usually produce more stable outputs, while top-p, top-k, min-p, and repeat penalty settings further control token selection and repetition. \`\`\`bash # Tune token selection and repetition behavior unsloth run \\ --model unsloth/Qwen3-1.7B-GGUF \\ --top-p 0.95 \\ --top-k 20 \\ --min-p 0.05 \\ --repeat-penalty 1.1 \`\`\` #### Increase context length and CPU threads Useful if you're working with large projects, long chats, or agent workflows that need more memory. \`\`\`bash # Use a larger context window and more CPU threads unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \\ -c 131072 \\ --threads 32 \`\`\` #### Expose the API on your local network By default, Unsloth only runs locally on your machine. You can expose the API to other devices on your network by binding to \`0.0.0.0\`. \`\`\`bash # Allow LAN devices to connect unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \\ -H 0.0.0.0 \\ -p 8888 \`\`\` #### Control reasoning behavior Some reasoning-capable models support additional flags for controlling thinking and reasoning behavior. \`\`\`bash # Disable reasoning / thinking output unsloth run \\ --model unsloth/Qwen3-1.7B-GGUF \\ --reasoning off \`\`\` \`\`\`bash # Enable reasoning mode unsloth run \\ --model unsloth/Qwen3-1.7B-GGUF \\ --reasoning on \`\`\` Reasoning support depends on the model and backend capabilities. #### Enable or disable server-side tools Control whether tools like web search and code execution are exposed by the inference server. \`\`\`bash # Explicitly enable tools unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \\ --enable-tools \`\`\` \`\`\`bash # Explicitly disable tools unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4\_K\_XL \\ --disable-tools \`\`\` Unsloth supports most llama-server runtime flags, including context sizing, GPU layers, threading, sampling, networking, and tool configuration. See the \[llama-server\](https://github.com/ggml-org/llama.cpp/tree/master/tools/server) documentation for the full list of supported runtime flags. #### \*\*Server-side tool policy\*\* \`unsloth run\` controls whether server-side tools (web search, code execution, etc.) are exposed by the inference server. Defaults are based on the bind address: \* \*\*\`127.0.0.1\` (localhost)\*\* — tools \*\*on\*\* by default. Only your machine can reach the server. \* \*\*\`0.0.0.0\` or any non-loopback address\*\* — tools \*\*off\*\* by default. A leaked API key on a network-exposed server means arbitrary code execution on the host. \*\*Flags:\*\* \* \`--enable-tools\` / \`--disable-tools\` — force on or off. On \`0.0.0.0\`, \`--enable-tools\` shows a y/N security prompt. \* \`--yes\` / \`-y\` — skip the prompt (for automation). The resolved policy is a process-level hard override — individual requests cannot bypass it via \`enable\_tools=true\` in the request body. ![](https://unsloth.ai/files/MaID7j6ybUcV10UN2VhO) \### 🌐 \*\*Endpoints\*\* Unsloth exposes these endpoints on whichever port it booted on (typically \`http://localhost:8000\` or \`http://localhost:8888\`): | Endpoint | Compatible with | Use it from | | --------------------------- | --------------------------- | --------------------------------------------------------------------- | | \`POST /v1/messages\` | Anthropic Messages API | Claude Code, Anthropic SDK, OpenClaw, anything that speaks Anthropic | | \`POST /v1/chat/completions\` | OpenAI Chat Completions API | OpenAI SDK, opencode, Cursor, Continue, Cline, Open WebUI, curl, etc. | | \`GET /v1/models\` | OpenAI models list | List the models currently loaded in Unsloth | Authenticate with an \`Authorization: Bearer sk-unsloth-…\` header on every request. {% hint style="info" %} You don't need to run different servers for the two formats. Unsloth handles both on the same port. {% endhint %} ### 🖇️ Connecting your client Unsloth enables you run local LLMs via most frameworks including \[Claude Code\](/docs/basics/claude-code.md), \[Codex\](/docs/basics/codex.md), \[OpenClaw\](/docs/integrations/openclaw.md), \[OpenCode\](/docs/integrations/opencode.md) and more. Click the specific tools below for a guide: {% columns %} {% column width="50%" %} {% content-ref url="/pages/w020xJgdCTBtTvfHtvye" %} \[Claude Code\](/docs/basics/claude-code.md) {% endcontent-ref %} {% content-ref url="/pages/PCjZ57h5pE0QccKyJMYD" %} \[OpenAI Codex\](/docs/basics/codex.md) {% endcontent-ref %} {% content-ref url="/pages/1cwX0SOqPoqLx7fQ2sIS" %} \[Curl & HTTP\](/docs/integrations/connect-curl-and-http-to-unsloth.md) {% endcontent-ref %} {% endcolumn %} {% column width="50%" %} {% content-ref url="/pages/CwQEpEmkKPmyEYdnEngt" %} \[OpenClaw\](/docs/integrations/openclaw.md) {% endcontent-ref %} {% content-ref url="/pages/qaA8ZjTxsH2GTuBOHyra" %} \[OpenCode\](/docs/integrations/opencode.md) {% endcontent-ref %} {% content-ref url="/pages/viZvzp58ObzZkXtCm0qv" %} \[Python SDK\](/docs/integrations/connect-python-sdk-to-unsloth.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} ### 🧰 Tool calling Both endpoints support function / tool calling in their native format, plus an Unsloth-specific shorthand for Unsloth's built-in tools. \*\*OpenAI-style tools:\*\* send \`tools\` and \`tool\_choice\` to \`/v1/chat/completions\` exactly as you would with OpenAI. Claude Code (via \`/v1/messages\`), opencode, Cursor, Continue, and Cline all work out of the box. \*\*Anthropic-style tools:\*\* send \`tools\` (with \`input\_schema\`) and \`tool\_choice\` to \`/v1/messages\` exactly as you would with Claude. Unsloth server side tools: Unsloth can execute Python, web search, and bash \*server-side\* and stream the results back as \`tool\_result\` events. Opt in by adding these extra fields to either endpoint: \`\`\`json { "messages": \[{"role": "user", "content": "What is 123 \* 456? Use Python."}\], "stream": true, "enable\_tools": true, "enabled\_tools": \["python", "web\_search","terminal"\], "session\_id": "my-session" } \`\`\` The model sees each tool's output on its next turn. For deeper coverage (schemas, streaming events, chaining), see . {% hint style="info" %} If you're using the Anthropic \`/v1/messages\` endpoint, \`tool\_choice\` maps cleanly: Anthropic \`auto\` → OpenAI \`auto\`, Anthropic \`any\` → OpenAI \`required\`, Anthropic \`{type: "tool", name: "x"}\` → OpenAI \`{type: "function", function: {name: "x"}}\`, Anthropic \`none\` → OpenAI \`none\`. {% endhint %} ### ❔ Troubleshooting \*\*\`401 Unauthorized\`\*\* : either the \`Authorization\` header is missing or the key is wrong. Keys must be passed as \`Authorization: Bearer sk-unsloth-…\`. If you lost the key, create a new one from \*\*Settings → API.\*\* Unsloth doesn't show old keys after creation. \*\*\`Lost connection to the model server\`\*\* : Unsloth couldn't reach the underlying llama.cpp server. Usually the model finished loading but crashed, or the model tab was closed inside Unsloth. Reload the model from \*\*New Chat\*\* and retry. \*\*Claude Code shows the default Anthropic model, not my local one\*\* : check all three env vars are exported in the \*\*same\*\* shell where you run \`claude\`: \`\`\`bash echo $ANTHROPIC\_BASE\_URL echo $ANTHROPIC\_AUTH\_TOKEN echo $ANTHROPIC\_MODEL \`\`\` Then run \`/model\` inside Claude Code to confirm. On Windows PowerShell use \`$env:ANTHROPIC\_BASE\_URL\` etc. \*\*\`stream: true\` returns a single JSON blob instead of SSE\*\* : make sure you're hitting the right path (\`/v1/messages\` or \`/v1/chat/completions\`) and that your HTTP client is actually consuming the response as a stream, not buffering it. \*\*I can't find the name of the model to add to opencode (or OpenClaw / any other client)\*\* : ask Unsloth directly. \`GET /v1/models\` returns the exact model ID you need to plug into the client's "Model ID" field: \`\`\`bash curl http://localhost:8888/v1/models \\ -H "Authorization: Bearer sk-unsloth-xxxxxxxxxxxx" \`\`\` You'll get back a JSON payload of the form \`{"data": \[{"id": "gemma-4-26B-A4B-it-GGUF", ...}\]}\`. Copy the \`id\` value, that's the string opencode's \*\*Model ID\*\* field (left column) and OpenClaw's \`models\[\].id\` expect. The display name on the right is whatever you want users to see. \*\*Tool calls aren't executed\*\* : The model needs to support tool calling for client-side tools (\`tools\` / \`tool\_choice\`). For Unsloth's built-in tools, remember to set \`enable\_tools: true\` \*\*and\*\* list the ones you want in \`enabled\_tools\` (e.g. \`\["python", "web\_search"\]\`). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/api.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/nouveau/studio/export.md). # Exporter des modèles avec Unsloth Studio Utiliser \[Unsloth Studio\](/docs/fr/nouveau/studio.md) pour exporter, enregistrer ou convertir des modèles en GGUF, Safetensors ou LoRA pour le déploiement, le partage ou l’inférence locale dans Unsloth, llama.cpp, Ollama, vLLM, et plus encore. Exportez un checkpoint entraîné ou convertissez n’importe quel modèle existant. ![](https://unsloth.ai/files/7f4d8c466857082cb4d586d66553185b401b1de1) {% stepper %} {% step %} ### Sélectionner l’exécution d’entraînement Commencez par sélectionner l’exécution d’entraînement à partir de laquelle vous souhaitez exporter. Chaque exécution représente une session d’entraînement complète et peut contenir plusieurs checkpoints. Après avoir choisi une exécution, sélectionnez le checkpoint à exporter. Un checkpoint est une version enregistrée du modèle créée pendant l’entraînement. ![](https://unsloth.ai/files/7eeae0aa1389d06edf94d9fbeac423deb27b5c73) {% endstep %} {% step %} ### Sélectionner le checkpoint Les checkpoints plus récents représentent généralement le modèle final entraîné, mais vous pouvez exporter n’importe quel checkpoint selon vos besoins. ![](https://unsloth.ai/files/3c088fdf13fb0417164d89c4c422318ff8955113) {% endstep %} {% step %} ### Méthodes d’exportation Selon votre flux de travail, vous pouvez exporter un modèle fusionné, les poids de l’adaptateur LoRA ou un modèle GGUF pour l’inférence locale. ![](https://unsloth.ai/files/709b49d77919cacb9fa7247defbcda531bcb641c) Chaque méthode d’exportation produit une version différente du modèle selon la façon dont vous prévoyez de l’exécuter ou de le partager. Le tableau ci-dessous explique ce que chaque option exporte. | Type d’exportation | Description | | ------------------ | --------------------------------------------------------------------------------------------------- | | Modèle fusionné | \*\*modèle 16 bits\*\* avec l’adaptateur LoRA fusionné dans les poids de base. | | LoRA uniquement | Exporte \*\*seulement les poids de l’adaptateur\*\*. Nécessite le modèle de base d’origine. | | GGUF / llama.cpp | Convertit le modèle en \*\*format GGUF\*\* pour Unsloth / llama.cpp \*\*/\*\* Ollama / LM Studio inférence. | | {% endstep %} | | {% step %} ### Exporter / Enregistrer localement Lors de l’exportation d’un modèle, vous pouvez choisir où les fichiers résultants doivent être enregistrés. Les modèles peuvent être téléchargés directement sur votre machine ou envoyés vers le Hugging Face Hub pour l’hébergement et le partage. Enregistrez les fichiers du modèle exporté directement sur votre machine. Cette option est utile pour exécuter le modèle localement, distribuer les fichiers manuellement ou l’intégrer à des outils d’inférence locaux. ![](https://unsloth.ai/files/5bb562a8e89589b2b5137ef15ddc45d37b9779bc) {% endstep %} {% step %} ### Publier sur le Hub Téléchargez le modèle exporté sur le Hugging Face Hub. Cela vous permet d’héberger, de partager et de déployer le modèle à partir d’un dépôt central. Vous aurez besoin d’un jeton d’écriture Hugging Face pour publier le modèle. ![](https://unsloth.ai/files/17fca4be2529601be8c805e046e4ef3324eb62ac) {% hint style="success" %} Si vous êtes déjà authentifié avec l’interface CLI Hugging Face, le jeton d’écriture peut être laissé vide. {% endhint %} {% endstep %} {% endstepper %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/nouveau/studio/export.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/models/qwen3.5/fine-tune.md). # Qwen3.5 Fine-tuning Guide You can now fine-tune \[Qwen3.5\](/docs/models/qwen3.5.md) model family (0.8B, 2B, 4B, 9B, 27B, 35B‑A3B, 122B‑A10B) with \[\*\*Unsloth\*\*\](https://github.com/unslothai/unsloth). Support includes both \[vision\](/docs/models/qwen3.5/fine-tune.md#vision-fine-tuning), text and \[RL\](#reinforcement-learning-rl) fine-tuning. \*\*Qwen3.5‑35B‑A3B\*\* - bf16 LoRA works on \*\*74GB VRAM.\*\* \* Unsloth makes Qwen3.5 train \*\*1.5× faster\*\* and uses \*\*50% less VRAM\*\* than FA2 setups. \* Qwen3.5 bf16 LoRA VRAM use: \*\*0.8B\*\*: 3GB • \*\*2B\*\*: 5GB • \*\*4B\*\*: 10GB • \*\*9B\*\*: 22GB • \*\*27B\*\*: 56GB \* Fine-tune \*\*0.8B\*\*, \*\*2B\*\* and \*\*4B\*\* bf16 LoRA via our \*\*free\*\* \*\*Google Colab notebooks\*\*: | \[Qwen3.5-\*\*0.8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) | \[Qwen3.5-4B \*\*GRPO\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision\_GRPO.ipynb) | | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | \* If you want to \*\*preserve reasoning\*\* ability, you can mix reasoning-style examples with direct answers (keep a minimum of 75% reasoning). Otherwise you can emit it fully. \* \*\*Full fine-tuning (FFT)\*\* works as well. Note it will use 4x more VRAM. \* Qwen3.5 is powerful for multilingual fine-tuning as it supports 201 languages. \* After fine-tuning, you can export to \[GGUF\](#saving-export-your-fine-tuned-model) (for llama.cpp/Ollama/etc.) or \[vLLM\](#saving-export-your-fine-tuned-model) \* \[Reinforcement Learning\](/docs/get-started/reinforcement-learning-rl-guide.md) (RL) for Qwen3.5 \[VLM RL\](/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md) also works via Unsloth inference. \* We have \*\*A100\*\* Colab notebooks for \[Qwen3.5‑27B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen\_3\_5\_27B\_A100\\(80GB\\).ipynb) and \[Qwen3.5‑35B‑A3B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_MoE.ipynb). If you’re on an older version (or fine-tuning locally), update first: {% columns %} {% column width="50%" %} Unsloth Studio: {% code expandable="true" %} \`\`\`bash unsloth studio update \`\`\` {% endcode %} {% endcolumn %} {% column width="50%" %} Unsloth code-based: \`\`\`bash pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` {% endcolumn %} {% endcolumns %} {% hint style="warning" %} \*\*Please use \`transformers v5\` for Qwen3.5. Older versions will not work. Unsloth automatically uses transformers v5 by default now (except for Colab environments).\*\* If training seems \*\*slower than usual\*\*, it’s because Qwen3.5 use custom Mamba Triton kernels. Compiling those kernels can take longer than normal, especially on T4 GPUs. It is not recommended to do QLoRA (4-bit) training on the Qwen3.5 models, no matter MoE or dense, due to higher than normal quantization differences. {% endhint %} ### MoE fine-tuning (35B, 122B) For MoE models like \*\*Qwen3.5‑35B‑A3B / 122B‑A10B / 397B‑A17B\*\*: \* You can use our \[Qwen3.5‑35B‑A3B (A100)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_MoE.ipynb) fine-tuning notebook \* Supports our recent \\~12x faster \[MoE training update\](/docs/basics/faster-moe.md) with >35% less VRAM & \\~6x longer context \* \*\*Best to use bf16 setups (e.g. LoRA or full fine-tuning)\*\* (MoE QLoRA 4‑bit is not recommended due to BitsandBytes limitations). \* Unsloth’s MoE kernels are enabled by default and can use different backends; you can switch with \`UNSLOTH\_MOE\_BACKEND\`. \* Router-layer fine-tuning is disabled by default for stability. \* Qwen3.5‑122B‑A10B - bf16 LoRA works on 256GB VRAM. If you're using multiGPUs, add \`device\_map = "balanced"\` or follow our \[multiGPU Guide\](/docs/basics/multi-gpu-training-with-unsloth.md). ### Quickstart #### 🦥 Unsloth Studio Guide Qwen3.5 can be run and fine-tuned in \[Unsloth Studio\](/docs/new/studio.md), our new open-source web UI for local AI. With Unsloth Studio, you can run models locally on \*\*MacOS, Windows\*\*, Linux and: {% columns %} {% column %} \* \[Train LLMs\](/docs/new/studio.md#no-code-training) 2x faster with 70% less VRAM \* Search, download, \[run GGUFs\](/docs/new/studio.md#run-models-locally) and safetensor models \* \[\*\*Self-healing\*\* tool calling\](/docs/new/studio.md#execute-code--heal-tool-calling) + \*\*web search\*\* \* \[\*\*Code execution\*\*\](/docs/new/studio.md#run-models-locally) (Python, Bash) \* \[Automatic inference\](/docs/new/studio.md#model-arena) parameter tuning (temp, top-p, etc.) \* Fast CPU + GPU inference via llama.cpp {% endcolumn %} {% column %} ![](https://unsloth.ai/files/92X2imx2aLf1xvLnVEb2) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Install Unsloth Run in your terminal: \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% hint style="success" %} \*\*Installation will be quick and take approx 1-2 mins.\*\* {% endhint %} {% endstep %} {% step %} #### Launch Unsloth \*\*MacOS, Linux, WSL and Windows:\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Then open \`http://localhost:8888\` in your browser.\*\* {% endstep %} {% step %} #### Train Qwen3.5 On first launch you will need to create a password to secure your account and sign in again later. You’ll then see a brief onboarding wizard to choose a model, dataset, and basic settings. You can skip it at any time. Search for Qwen3.5 in the search bar and select your desired model and dataset. Next, adjust your hyperparameters, context length as desired. ![](https://unsloth.ai/files/8vzEvUl3LaefJAu9xfFx) {% endstep %} {% step %} #### Monitor training progress After you click start training, you will be able to monitor and observe the training progress of the model. The training loss should be steadily decreasing.\\ Once done, the model will be automatically saved. ![](https://unsloth.ai/files/JBoSp5HwP0aZgI7vpguh) {% endstep %} {% step %} #### Export your fine-tuned model Once done, Unsloth Studio allows you to export the model to GGUF, safetensor etc formats. ![](https://unsloth.ai/files/Bz6jO9RYzYHWW7sFB4mA) {% endstep %} {% endstepper %} #### Unsloth Core (code-based) guide: Below is a minimal SFT recipe (works for “text-only” fine-tuning). See also our \[vision fine-tuning\](/docs/basics/vision-fine-tuning.md) section. {% hint style="info" %} Qwen3.5 is “Causal Language Model with Vision Encoder” (it’s a unified VLM), so ensure you have the usual vision deps installed (\`torchvision\`, \`pillow\`) if needed, and keep Transformers up-to-date. Use the latest Transformers for Qwen3.5. \*\*If you'd like to do\*\* \[\*\*GRPO\*\*\](/docs/get-started/reinforcement-learning-rl-guide.md)\*\*, it works in Unsloth if you disable fast vLLM inference and use Unsloth inference instead. Follow our\*\* \[\*\*Vision RL\*\*\](/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md) \*\*notebook examples.\*\* {% endhint %} {% code expandable="true" %} \`\`\`python from unsloth import FastLanguageModel import torch from datasets import load\_dataset from trl import SFTTrainer, SFTConfig max\_seq\_length = 2048 # start small; scale up after it works # Example dataset (replace with yours). Needs a "text" column. url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified\_chip2.jsonl" dataset = load\_dataset("json", data\_files={"train": url}, split="train") model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "Qwen/Qwen3.5-27B", max\_seq\_length = max\_seq\_length, load\_in\_4bit = False, # MoE QLoRA not recommended, dense 27B is fine load\_in\_16bit = True, # bf16/16-bit LoRA full\_finetuning = False, ) model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \[ "q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj", \], lora\_alpha = 16, lora\_dropout = 0, bias = "none", # "unsloth" checkpointing is intended for very long context + lower VRAM use\_gradient\_checkpointing = "unsloth", random\_state = 3407, max\_seq\_length = max\_seq\_length, ) trainer = SFTTrainer( model = model, train\_dataset = dataset, tokenizer = tokenizer, args = SFTConfig( max\_seq\_length = max\_seq\_length, per\_device\_train\_batch\_size = 1, gradient\_accumulation\_steps = 4, warmup\_steps = 10, max\_steps = 100, logging\_steps = 1, output\_dir = "outputs\_qwen35", optim = "adamw\_8bit", seed = 3407, dataset\_num\_proc = 1, ), ) trainer.train() \`\`\` {% endcode %} {% hint style="info" %} If you OOM: \* Drop \`per\_device\_train\_batch\_size\` to \*\*1\*\* and/or reduce \`max\_seq\_length\`. \* Keep \`use\_\`\[\`gradient\_checkpointing\`\](/docs/blog/500k-context-length-fine-tuning.md#unsloth-gradient-checkpointing-enhancements)\`="unsloth"\` on (it’s designed to reduce VRAM use and extend context length). {% endhint %} \*\*Loader example for MoE (bf16 LoRA):\*\* \`\`\`python import os import torch from unsloth import FastModel model, tokenizer = FastModel.from\_pretrained( model\_name = "unsloth/Qwen3.5-35B-A3B", max\_seq\_length = 2048, load\_in\_4bit = False, # MoE QLoRA not recommended, dense 27B is fine load\_in\_16bit = True, # bf16/16-bit LoRA full\_finetuning = False, ) \`\`\` Once loaded, you’ll attach LoRA adapters and train similarly to the SFT example above. ### Vision fine-tuning Unsloth supports \[vision fine-tuning\](/docs/basics/vision-fine-tuning.md) for the multimodal Qwen3.5 models. Use the below Qwen3.5 notebooks and change the respective model names to your desired Qwen3.5 model. | \[Qwen3.5-\*\*0.8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) | Qwen3.5-\*\*9B\*\* | | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | -------------- | \* \[Qwen3-VL GRPO/GSPO RL notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) (change model name to Qwen3.5-4B etc.) \*\*Disabling Vision / Text-only fine-tuning:\*\* To fine-tune vision models, we now allow you to select which parts of the mode to finetune. You can select to only fine-tune the vision layers, or the language layers, or the attention / MLP layers! We set them all on by default! {% code expandable="true" %} \`\`\`python model = FastVisionModel.get\_peft\_model( model, finetune\_vision\_layers = True, # False if not finetuning vision layers finetune\_language\_layers = True, # False if not finetuning language layers finetune\_attention\_modules = True, # False if not finetuning attention layers finetune\_mlp\_modules = True, # False if not finetuning MLP layers r = 16, # The larger, the higher the accuracy, but might overfit lora\_alpha = 16, # Recommended alpha == r at least lora\_dropout = 0, bias = "none", random\_state = 3407, use\_rslora = False, # We support rank stabilized LoRA loftq\_config = None, # And LoftQ target\_modules = "all-linear", # Optional now! Can specify a list if needed modules\_to\_save=\[ "lm\_head", "embed\_tokens", \], ) \`\`\` {% endcode %} In order to fine-tune or train Qwen3.5 with multi-images, view our \[\*\*multi-image vision guide\*\*\](/docs/basics/vision-fine-tuning.md#multi-image-training)\*\*.\*\* ### Reinforcement Learning (RL) You can now train Qwen3.5 with RL, GSPO, GRPO etc with \[our free notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision\_GRPO.ipynb): {% embed url="" %} You can run Qwen3.5 RL with Unsloth even though it is not supported by vLLM, by setting \`fast\_inference=False\` when loading the model: \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( model\_name="unsloth/Qwen3.5-4B", fast\_inference=False, ) \`\`\` ### Saving / export fine-tuned model You can view our specific inference / deployment guides for \[Unsloth Studio\](/docs/new/studio/export.md), \[llama.cpp\](/docs/basics/inference-and-deployment/saving-to-gguf.md), \[vLLM\](/docs/basics/inference-and-deployment/vllm-guide.md), \[llama-server\](/docs/basics/inference-and-deployment/llama-server-and-openai-endpoint.md), \[Ollama\](/docs/basics/inference-and-deployment/saving-to-ollama.md). #### Save to GGUF Unsloth supports saving directly to GGUF: \`\`\`python model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q4\_k\_m") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q8\_0") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "f16") \`\`\` Or push GGUFs to Hugging Face: \`\`\`python model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q4\_k\_m") model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q8\_0") \`\`\` If the exported model behaves worse in another runtime, Unsloth flags the most common cause: \*\*wrong chat template / EOS token at inference time\*\* (you must use the same chat template you trained with). #### Save to vLLM {% hint style="warning" %} vLLM version \`0.16.0\` does not support Qwen3.5. Wait until \`0.170\` or try the Nightly release. {% endhint %} To save to 16-bit for vLLM, use: {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("finetuned\_model", tokenizer, save\_method = "merged\_16bit") ## OR to upload to HuggingFace: model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "merged\_16bit", token = "") \`\`\` {% endcode %} To save just the LoRA adapters, either use: \`\`\`python model.save\_pretrained("finetuned\_lora") tokenizer.save\_pretrained("finetuned\_lora") \`\`\` Or use our builtin function: {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("finetuned\_model", tokenizer, save\_method = "lora") ## OR to upload to HuggingFace model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "lora", token = "") \`\`\` {% endcode %} For more details read our inference guides: {% columns %} {% column width="50%" %} {% content-ref url="/pages/gEugERiAw2ztDNt98JVR" %} \[Inference & Deployment\](/docs/basics/inference-and-deployment.md) {% endcontent-ref %} {% content-ref url="/pages/T7ZPf3SNAwDykZNgXptE" %} \[GGUF & llama.cpp\](/docs/basics/inference-and-deployment/saving-to-gguf.md) {% endcontent-ref %} {% endcolumn %} {% column width="50%" %} {% content-ref url="/pages/5ZU2kPF2eJ7VK0GeEUhu" %} \[Model Export\](/docs/new/studio/export.md) {% endcontent-ref %} {% content-ref url="/pages/fhJtaLFFXVsGnbMUiACo" %} \[vLLM\](/docs/basics/inference-and-deployment/vllm-guide.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/models/qwen3.5/fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf.md). # Saving to GGUF Saving models to 16bit for GGUF so you can use it for \[Unsloth Studio\](/docs/new/studio.md), Ollama, llama.cpp and more! {% tabs %} {% tab title="Locally" %} To save to GGUF, use the below to save locally: \`\`\`python model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q4\_k\_m") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q8\_0") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "f16") \`\`\` To push to Hugging Face hub: \`\`\`python model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q4\_k\_m") model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q8\_0") \`\`\` All supported quantization options for \`quantization\_method\` are listed below: \`\`\`python # https://github.com/ggml-org/llama.cpp/blob/master/examples/quantize/quantize.cpp#L19 ALLOWED\_QUANTS = \\ { "not\_quantized" : "Recommended. Fast conversion. Slow inference, big files.", "fast\_quantized" : "Recommended. Fast conversion. OK inference, OK file size.", "quantized" : "Recommended. Slow conversion. Fast inference, small files.", "f32" : "Not recommended. Retains 100% accuracy, but super slow and memory hungry.", "f16" : "Fastest conversion + retains 100% accuracy. Slow and memory hungry.", "q8\_0" : "Fast conversion. High resource use, but generally acceptable.", "q4\_k\_m" : "Recommended. Uses Q6\_K for half of the attention.wv and feed\_forward.w2 tensors, else Q4\_K", "q5\_k\_m" : "Recommended. Uses Q6\_K for half of the attention.wv and feed\_forward.w2 tensors, else Q5\_K", "q2\_k" : "Uses Q4\_K for the attention.wv and feed\_forward.w2 tensors, Q2\_K for the other tensors.", "q3\_k\_l" : "Uses Q5\_K for the attention.wv, attention.wo, and feed\_forward.w2 tensors, else Q3\_K", "q3\_k\_m" : "Uses Q4\_K for the attention.wv, attention.wo, and feed\_forward.w2 tensors, else Q3\_K", "q3\_k\_s" : "Uses Q3\_K for all tensors", "q4\_0" : "Original quant method, 4-bit.", "q4\_1" : "Higher accuracy than q4\_0 but not as high as q5\_0. However has quicker inference than q5 models.", "q4\_k\_s" : "Uses Q4\_K for all tensors", "q4\_k" : "alias for q4\_k\_m", "q5\_k" : "alias for q5\_k\_m", "q5\_0" : "Higher accuracy, higher resource usage and slower inference.", "q5\_1" : "Even higher accuracy, resource usage and slower inference.", "q5\_k\_s" : "Uses Q5\_K for all tensors", "q6\_k" : "Uses Q8\_K for all tensors", "iq2\_xxs" : "2.06 bpw quantization", "iq2\_xs" : "2.31 bpw quantization", "iq3\_xxs" : "3.06 bpw quantization", "q3\_k\_xs" : "3-bit extra small quantization", } \`\`\` {% endtab %} {% tab title="Manual Saving" %} First save your model to 16bit: \`\`\`python model.save\_pretrained\_merged("merged\_model", tokenizer, save\_method = "merged\_16bit",) \`\`\` Then use the terminal and do: {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp python llama.cpp/convert\_hf\_to\_gguf.py FOLDER --outfile OUTPUT --outtype f16 \`\`\` {% endcode %} Or follow the steps at using the model name "merged\\\_model" to merge to GGUF. {% endtab %} {% endtabs %} ### Running in Unsloth works well, but after exporting & running on other platforms, the results are poor You might sometimes encounter an issue where your model runs and produces good results on Unsloth, but when you use it on another platform like Ollama or vLLM, the results are poor or you might get gibberish, endless/infinite generations \*or\* repeated outputs\*\*.\*\* \* The most common cause of this error is using an \*\*incorrect chat template\*\*\*\*.\*\* It’s essential to use the SAME chat template that was used when training the model in Unsloth and later when you run it in another framework, such as llama.cpp or Ollama. When inferencing from a saved model, it's crucial to apply the correct template. \* You must use the correct \`eos token\`. If not, you might get gibberish on longer generations. \* It might also be because your inference engine adds an unnecessary "start of sequence" token (or the lack of thereof on the contrary) so ensure you check both hypotheses! \* \*\*Use our conversational notebooks to force the chat template - this will fix most issues.\*\* \* Qwen-3 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* Gemma-3 4B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) \* Llama-3.2 3B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* Phi-4 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* Mistral v0.3 7B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \*\*More notebooks in our\*\* \[\*\*notebooks docs\*\*\](/docs/get-started/unsloth-notebooks.md) ### Saving to GGUF / vLLM 16bit crashes You can try reducing the maximum GPU usage during saving by changing \`maximum\_memory\_usage\`. The default is \`model.save\_pretrained(..., maximum\_memory\_usage = 0.75)\`. Reduce it to say 0.5 to use 50% of GPU peak memory or lower. This can reduce OOM crashes during saving. ### How do I manually save to GGUF? First save your model to 16bit via: {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("merged\_model", tokenizer, save\_method = "merged\_16bit",) \`\`\` {% endcode %} Compile llama.cpp from source like below: {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} Then, save the model to F16: \`\`\`bash python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-F16.gguf --outtype f16 \\ --split-max-size 50G \`\`\` \`\`\`bash # For BF16: python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-BF16.gguf --outtype bf16 \\ --split-max-size 50G # For Q8\_0: python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-Q8\_0.gguf --outtype q8\_0 \\ --split-max-size 50G \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide.md). # Guide des hyperparamètres de fine-tuning LoRA Les hyperparamètres LoRA sont des réglages ajustables qui régissent la façon dont l'Adaptation à Faible Rang \[affine\](/docs/fr/commencer/fine-tuning-llms-guide.md) les LLM. Avec de nombreux choix (par ex., taux d'apprentissage et époques) et d'innombrables combinaisons, choisir les bonnes valeurs est essentiel pour la précision, la stabilité, la qualité et la réduction des hallucinations. Bien fait, \*\*LoRA peut égaler les performances d'un fine-tuning complet\*\* tout en utilisant 4× moins de VRAM. Vous apprendrez les meilleures pratiques pour ces paramètres, basées sur des enseignements tirés de centaines d'articles de recherche et d'expériences, et verrez comment ils impactent le modèle. \*\*Bien que nous recommandions d'utiliser les paramètres par défaut d'Unsloth\*\*, comprendre ces concepts vous donnera un contrôle total.\\ \\ L'objectif est de modifier les valeurs des hyperparamètres pour augmenter la précision tout en contrebalançant \[\*\*le surapprentissage ou le sous-apprentissage\*\*\](#overfitting-poor-generalization-too-specialized). Le surapprentissage se produit lorsque le modèle mémorise les données d'entraînement, nuisant à sa capacité à généraliser à de nouvelles entrées non vues. L'objectif est d'obtenir un modèle qui généralise bien, pas un modèle qui se contente de mémoriser. {% columns %} {% column %} #### :question:Mais qu'est-ce que LoRA ? Dans les LLM, nous avons des poids de modèle. Llama 70B a 70 milliards de nombres. Plutôt que de changer les 70 milliards de nombres, nous ajoutons des matrices minces A et B à chaque poids, et optimisons celles-ci. Cela signifie que nous n'optimisons que 1 % des poids. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/5bc587b3ea8dfe36f7bd2869657ce08eeb810a94) Au lieu d'optimiser les poids du modèle (en jaune), nous optimisons 2 matrices minces A et B. {% endcolumn %} {% endcolumns %} ## :1234: Hyperparamètres clés du Fine-tuning ### \*\*Taux d'apprentissage\*\* Définit dans quelle mesure les poids du modèle sont ajustés à chaque étape d'entraînement. \* \*\*Taux d'apprentissage élevés\*\* : Mènent à une convergence initiale plus rapide mais peuvent rendre l'entraînement instable ou empêcher de trouver un minimum optimal si trop élevés. \* \*\*Taux d'apprentissage faibles\*\* : Donnent un entraînement plus stable et précis mais peuvent nécessiter plus d'époques pour converger, augmentant le temps total d'entraînement. Bien que l'on pense souvent que des taux faibles provoquent du sous-apprentissage, ils peuvent en réalité entraîner \*\*le surapprentissage\*\* ou même empêcher le modèle d'apprendre. \* \*\*Plage typique\*\*: \`2e-4\` (0.0002) à \`5e-6\` (0.000005).\\ :green\\\_square: \*\*\*Pour un fine-tuning LoRA/QLoRA normal\*\*\*, \*nous recommandons\* \*\*\`2e-4\`\*\* \*comme point de départ.\*\\ :blue\\\_square: \*\*\*Pour l'apprentissage par renforcement\*\* (DPO, GRPO etc.), nous recommandons\* \*\*\`5e-6\` .\*\*\\ :white\\\_large\\\_square: \*\*\*Pour le fine-tuning complet,\*\* des taux d'apprentissage plus faibles sont généralement plus appropriés.\* ### \*\*Époques\*\* Le nombre de fois que le modèle voit l'ensemble complet de données d'entraînement. \* \*\*Plus d'époques :\*\* Peuvent aider le modèle à mieux apprendre, mais un grand nombre peut le faire \*\*mémoriser les données d'entraînement\*\*, nuisant à sa performance sur de nouvelles tâches. \* \*\*Moins d'époques :\*\* Réduisent le temps d'entraînement et peuvent prévenir le surapprentissage, mais peuvent aboutir à un modèle sous-entraîné si le nombre est insuffisant pour que le modèle apprenne les motifs sous-jacents du jeu de données. \* \*\*Recommandé :\*\* 1-3 époques. Pour la plupart des jeux de données basés sur des instructions, s'entraîner plus de 3 époques offre des rendements décroissants et augmente le risque de surapprentissage. ### \*\*LoRA ou QLoRA\*\* LoRA utilise une précision 16 bits, tandis que QLoRA est une méthode de fine-tuning en 4 bits. \* \*\*LoRA :\*\* Fine-tuning 16 bits. C'est légèrement plus rapide et légèrement plus précis, mais consomme beaucoup plus de VRAM (4× plus que QLoRA). Recommandé pour les environnements 16 bits et les scénarios où la précision maximale est requise. \* \*\*QLoRA :\*\* Fine-tuning 4 bits. Légèrement plus lent et marginalement moins précis, mais utilise beaucoup moins de VRAM (4× moins).\\ :sloth: \*LLaMA 70B tient dans <48GB de VRAM avec QLoRA dans Unsloth -\* \[\*plus de détails ici\*\](https://unsloth.ai/blog/llama3-3)\*.\* ### Hyperparamètres et recommandations : | Hyperparamètre | Fonction | Paramètres recommandés | | --- | --- | --- | | **Rang LoRA** (`r`) | Contrôle le nombre de paramètres entraînables dans les matrices adaptatrices LoRA. Un rang plus élevé augmente la capacité du modèle mais aussi l'utilisation mémoire. | 8, 16, 32, 64, 128

Choisissez 16 ou 32 | | **Alpha LoRA** (`lora_alpha`) | Ajuste la force des ajustements du fine-tuning par rapport au rang (`r`). | `r` (standard) ou `r * 2` (heuristique courante). [Plus de détails ici](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide.md#lora-alpha-and-rank-relationship)
. | | **Dropout LoRA** | Une technique de régularisation qui met aléatoirement une fraction des activations LoRA à zéro pendant l'entraînement pour prévenir le surapprentissage. **Pas très utile**, donc nous le réglons par défaut à 0. | 0 (par défaut) à 0.1 | | **Décroissance des poids** | Un terme de régularisation qui pénalise les poids importants pour prévenir le surapprentissage et améliorer la généralisation. N'utilisez pas des valeurs trop élevées ! | 0.01 (recommandé) - 0.1 | | **Étapes d'échauffement** | Augmente progressivement le taux d'apprentissage au début de l'entraînement. | 5-10% du nombre total d'étapes | | **Type de scheduler** | Ajuste dynamiquement le taux d'apprentissage pendant l'entraînement. | `linéaire` ou `cosine` | | **Graine (`random_state`)** | Un nombre fixe pour assurer la reproductibilité des résultats. | N'importe quel entier (par ex., `42`, `3407`) | | **Modules cibles** | Spécifiez quelles parties du modèle vous souhaitez appliquer les adaptateurs LoRA — soit l'attention, le MLP, ou les deux.


Attention : `q_proj, k_proj, v_proj, o_proj`

MLP : `gate_proj, up_proj, down_proj` | Recommandé de cibler toutes les principales couches linéaires : `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj`. | \## :deciduous\\\_tree: Équivalence entre accumulation de gradients et taille de lot ### Taille de lot effective Configurer correctement votre taille de lot est crucial pour équilibrer la stabilité de l'entraînement avec les limitations de VRAM de votre GPU. Ceci est géré par deux paramètres dont le produit est la \*\*Taille de lot effective\*\*.\\ \\ \*\*Taille de lot effective\*\* = \`batch\_size \* gradient\_accumulation\_steps\` \* Un \*\*Taille de lot effective plus grande\*\* conduit généralement à un entraînement plus fluide et plus stable. \* Un \*\*Taille de lot effective plus petite\*\* peut introduire davantage de variance. Bien que chaque tâche soit différente, la configuration suivante offre un excellent point de départ pour obtenir une \*\*Taille de lot effective\*\* de 16, qui fonctionne bien pour la plupart des tâches de fine-tuning sur les GPU modernes. | Paramètre | Description | Réglage recommandé | | ------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- | | \*\*Taille de lot\*\* (\`batch\_size\`) | Le nombre d'exemples traités dans une seule passe avant/arrière sur un GPU. **Principal facteur d'utilisation de la VRAM**. Des valeurs plus élevées peuvent améliorer l'utilisation du matériel et accélérer l'entraînement, mais seulement si elles tiennent en mémoire. | 2 | | \*\*Accumulation de gradients\*\* (\`gradient\_accumulation\_steps\`) | Le nombre de micro-lots à traiter avant d'effectuer une seule mise à jour des poids du modèle. **Principal facteur du temps d'entraînement.** Permet de simuler une plus grande `batch\_size` pour économiser de la VRAM. Des valeurs plus élevées augmentent le temps d'entraînement par époque. | 8 | | \*\*Taille de lot effective\*\* (Calculé) | La véritable taille de lot utilisée pour chaque mise à jour de gradient. Elle influence directement la stabilité de l'entraînement, la qualité et la performance finale du modèle. | 4 à 16 Recommandé : 16 (à partir de 2 \\\* 8) | ### Le compromis VRAM & performance Supposez que vous souhaitiez 32 échantillons de données par étape d'entraînement. Vous pouvez alors utiliser l'une des configurations suivantes : \* \`batch\_size = 32, gradient\_accumulation\_steps = 1\` \* \`batch\_size = 16, gradient\_accumulation\_steps = 2\` \* \`batch\_size = 8, gradient\_accumulation\_steps = 4\` \* \`batch\_size = 4, gradient\_accumulation\_steps = 8\` \* \`batch\_size = 2, gradient\_accumulation\_steps = 16\` \* \`batch\_size = 1, gradient\_accumulation\_steps = 32\` Bien que toutes ces configurations soient équivalentes pour les mises à jour des poids du modèle, elles ont des exigences matérielles très différentes. La première configuration (\`batch\_size = 32\`) utilise le \*\*plus de VRAM\*\* et échouera probablement sur la plupart des GPU. La dernière configuration (\`batch\_size = 1\`) utilise le \*\*le moins de VRAM,\*\* mais au prix d'un entraînement légèrement plus lent\*\*.\*\* Pour éviter les erreurs OOM (out of memory), préférez toujours définir un \`batch\_size\` plus petit et augmentez \`gradient\_accumulation\_steps\` pour atteindre votre \*\*Taille de lot effective\*\*. ### :sloth: Correction d'accumulation de gradient d'Unsloth L'accumulation de gradients et les tailles de lot \*\*sont désormais entièrement équivalentes dans Unsloth\*\* grâce à nos corrections de bugs pour l'accumulation de gradients. Nous avons implémenté des corrections spécifiques qui résolvent un problème courant où les deux méthodes ne produisaient pas les mêmes résultats. C'était un défi connu dans la communauté, mais pour les utilisateurs d'Unsloth, les deux méthodes sont désormais interchangeables. \[Lisez notre article de blog\](https://unsloth.ai/blog/gradient) pour plus de détails. Avant nos corrections, des combinaisons de \`batch\_size\` et \`gradient\_accumulation\_steps\` qui donnaient le même \*\*Taille de lot effective\*\* (c.-à-d., \`batch\_size × gradient\_accumulation\_steps = 16\`) n'entraînaient pas un comportement d'entraînement équivalent. Par exemple, des configurations comme \`b1/g16\`, \`b2/g8\`, \`b4/g4\`, \`b8/g2\`, et \`b16/g1\` ont toutes un \*\*Taille de lot effective\*\* de 16, mais comme le montre le graphe, les courbes de perte ne s'alignaient pas lors de l'utilisation de l'accumulation de gradient standard : ![](https://unsloth.ai/files/42213b383bc6da53153950c2608c397b3fd3a581) (Avant - Accumulation de gradient standard) Après l'application de nos corrections, les courbes de perte s'alignent désormais correctement, indépendamment de la façon dont le \*\*Taille de lot effective\*\* de 16 est atteint : ![](https://unsloth.ai/files/e839528f5f698a4e0399f509af9d62a40ed5cb9f) (Après - 🦥 Accumulation de gradient Unsloth) \## 🦥 \*\*Hyperparamètres LoRA dans Unsloth\*\* Ce qui suit démontre une configuration standard. \*\*Bien qu'Unsloth fournisse des valeurs par défaut optimisées\*\*, comprendre ces paramètres est essentiel pour un réglage manuel. ![](https://unsloth.ai/files/b1707d5525dacc5e699f42ba5cad119dbc8e9d00) 1\. \`\`\`python r = 16, # Choisissez n'importe quel nombre > 0 ! Suggestion 8, 16, 32, 64, 128 \`\`\` Le rang (\`r\`) du processus de fine-tuning. Un rang plus élevé utilise plus de mémoire et sera plus lent, mais peut augmenter la précision sur des tâches complexes. Nous suggérons des rangs comme 8 ou 16 (pour des fine-tunings rapides) et jusqu'à 128. Utiliser un rang trop grand peut provoquer du surapprentissage et nuire à la qualité de votre modèle.\\\\ 2. \`\`\`python target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\], \`\`\` Pour des performances optimales, \*\*LoRA devrait être appliqué à toutes les principales couches linéaires\*\*. \[La recherche a montré\](#lora-target-modules-and-qlora-vs-lora) que cibler toutes les couches majeures est crucial pour égaler les performances d'un fine-tuning complet. Bien qu'il soit possible de retirer des modules pour réduire l'utilisation mémoire, nous le déconseillons fortement afin de préserver la qualité maximale, car les économies sont minimes.\\\\ 3. \`\`\`python lora\_alpha = 16, \`\`\` Un facteur d'échelle qui contrôle la force des ajustements du fine-tuning. Le régler égal au rang (\`r\`) est une base fiable. Une heuristique populaire et efficace est de le régler au double du rang (\`r \* 2\`), ce qui pousse le modèle à apprendre de manière plus agressive en donnant plus de poids aux mises à jour LoRA. \[Plus de détails ici\](#lora-alpha-and-rank-relationship).\\\\ 4. \`\`\`python lora\_dropout = 0, # Accepte n'importe quelle valeur, mais = 0 est optimisé \`\`\` Une technique de régularisation qui aide à \[prévenir le surapprentissage\](#overfitting-poor-generalization-too-specialized) en mettant aléatoirement une fraction des activations LoRA à zéro à chaque étape d'entraînement. \[Des recherches récentes suggèrent\](https://arxiv.org/abs/2410.09692) que pour \*\*les courts entraînements\*\* courants dans le fine-tuning, \`lora\_dropout\` peut être un régularisateur peu fiable.\\ 🦥 \*Le code interne d'Unsloth peut optimiser l'entraînement lorsque\* \`lora\_dropout = 0\`\*, le rendant légèrement plus rapide, mais nous recommandons une valeur non nulle si vous suspectez du surapprentissage.\*\\\\ 5. \`\`\`python bias = "none", # Accepte n'importe quelle valeur, mais = "none" est optimisé \`\`\` Laissez ceci sur \`"none"\` pour un entraînement plus rapide et une utilisation mémoire réduite. Ce réglage évite d'entraîner les termes de biais dans les couches linéaires, ce qui ajoute des paramètres entraînables pour peu ou pas de gain pratique.\\\\ 6. \`\`\`python use\_gradient\_checkpointing = "unsloth", # True ou "unsloth" pour un contexte très long \`\`\` Les options sont \`True\`, \`False\`, et \`"unsloth"\`.\\ 🦥 \*Nous recommandons\* \`"unsloth"\` \*car cela réduit l'utilisation mémoire d'environ 30% supplémentaires et prend en charge des fine-tunings avec un contexte extrêmement long. Vous pouvez en lire plus sur\* \[\*notre article de blog sur l'entraînement avec long contexte\*\](https://unsloth.ai/blog/long-context)\*.\*\\\\ 7. \`\`\`python random\_state = 3407, \`\`\` La graine pour assurer des exécutions déterministes et reproductibles. L'entraînement implique des nombres aléatoires, donc fixer une graine est essentiel pour des expériences cohérentes.\\\\ 8. \`\`\`python use\_rslora = False, # Nous supportons le LoRA stabilisé par le rang \`\`\` Une fonctionnalité avancée qui implémente \[\*\*Rank-Stabilized LoRA\*\*\](https://arxiv.org/abs/2312.03732). Si réglé sur \`True\`, le facteur d'échelle effectif devient \`lora\_alpha / sqrt(r)\` au lieu du standard \`lora\_alpha / r\`. Cela peut parfois améliorer la stabilité, en particulier pour des rangs élevés. \[Plus de détails ici\](#lora-alpha-and-rank-relationship).\\\\ 9. \`\`\`python loftq\_config = None, # Et LoftQ \`\`\` Une technique avancée, proposée dans \[\*\*LoftQ\*\*\](https://arxiv.org/abs/2310.08659), initialise les matrices LoRA avec les 'r' principaux vecteurs singuliers des poids pré-entraînés. Cela peut améliorer la précision mais provoquer un pic significatif de mémoire au démarrage de l'entraînement. ### \*\*Vérification des mises à jour des poids LoRA :\*\* Lors de la validation que \*\*LoRA\*\* les poids de l'adaptateur ont été mis à jour après le fine-tuning, évitez d'utiliser \*\*np.allclose()\*\* pour la comparaison. Cette méthode peut manquer des changements subtils mais significatifs, en particulier dans \*\*LoRA A\*\*, qui est initialisée avec de petites valeurs gaussiennes. Ces changements peuvent ne pas être considérés comme significatifs sous des tolérances numériques larges. Merci aux \[contributeurs\](https://github.com/unslothai/unsloth/issues/3035) pour cette section. Pour confirmer de manière fiable les mises à jour des poids, nous recommandons : \* Utiliser \*\*des comparaisons de sommes de contrôle ou de hachage\*\* (par ex., MD5) \* Calculer la \*\*somme des différences absolues\*\* entre tenseurs \* Inspecter les statistiques de\*\*tenseur\*\* (par ex., moyenne, variance) manuellement \* Ou utiliser \*\*np.array\\\_equal()\*\* si une égalité exacte est attendue ## :triangular\\\_ruler:Relation entre Alpha LoRA et Rang {% hint style="success" %} Il est préférable de définir \`lora\_alpha = 2 \* lora\_rank\` ou \`lora\_alpha = lora\_rank\` {% endhint %} {% columns %} {% column width="50%" %} $$ \\hat{W} = W + \\frac{\\alpha}{\\text{rank}} \\times AB $$ ![](https://unsloth.ai/files/909fcdd38704d8f954ef3845489ded63c1650052) rsLoRA autres options d'échelle. sqrt(r) est la meilleure. $$ \\hat{W}\\\_{\\text{rslora}} = W + \\frac{\\alpha}{\\sqrt{\\text{rank}}} \\times AB $$ {% endcolumn %} {% column %} La formule pour LoRA est à gauche. Nous devons mettre à l'échelle les matrices minces A et B par alpha divisé par le rang. \*\*Cela signifie que nous devons garder alpha/rang au moins = 1\*\*. Selon le \[article rsLoRA (rank stabilized lora)\](https://arxiv.org/abs/2312.03732), nous devrions plutôt mettre à l'échelle alpha par la racine carrée du rang. D'autres options existent, mais théoriquement c'est l'optimum. Le graphique de gauche montre d'autres rangs et leurs perplexités (plus bas est meilleur). Pour activer cela, réglez \`use\_rslora = True\` dans Unsloth. Notre recommandation est de régler \*\*alpha égal au rang, ou au moins 2 fois le rang.\*\* Cela signifie alpha/rang = 1 ou 2. {% endcolumn %} {% endcolumns %} ## :dart: Modules cibles LoRA et QLoRA vs LoRA {% hint style="success" %} Utilisez :\\ \`target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\]\` pour cibler à la fois \*\*MLP\*\* et \*\*attention\*\* couches pour augmenter la précision. \*\*QLoRA utilise une précision 4 bits\*\*, réduisant l'utilisation de la VRAM de plus de 75%. \*\*LoRA (16 bits)\*\* est légèrement plus précis et plus rapide. {% endhint %} Selon des expériences empiriques et des articles de recherche comme l'original \[article QLoRA\](https://arxiv.org/pdf/2305.14314), il est préférable d'appliquer LoRA à la fois aux couches d'attention et MLP. {% columns %} {% column %} ![](https://unsloth.ai/files/d885d0af57659cf586a41f2e83267fe1acf93b34) {% endcolumn %} {% column %} Le graphique montre les scores RougeL (plus élevé est mieux) pour différentes configurations de modules cibles, comparant LoRA et QLoRA. Les 3 premiers points montrent : 1. \*\*QLoRA-All :\*\* LoRA appliqué à toutes les couches FFN/MLP et Attention.\\ :fire: \*Celle-ci donne les meilleures performances globales.\* 2. \*\*QLoRA-FFN\*\* : LoRA uniquement sur FFN.\\ Équivalent à : \`gate\_proj\`, \`up\_proj\`, \`down\_proj.\` 3. \*\*QLoRA-Attention\*\* : LoRA appliqué uniquement aux couches d'Attention.\\ Équivalent à : \`q\_proj\`, \`k\_proj\`, \`v\_proj\`, \`o\_proj\`. {% endcolumn %} {% endcolumns %} ## :sunglasses: S'entraîner uniquement sur les complétions, en masquant les entrées Le \[article QLoRA\](https://arxiv.org/pdf/2305.14314) montre que masquer les entrées et \*\*s'entraîner seulement sur les complétions\*\* (sorties ou messages de l'assistant) peut encore \*\*augmenter la précision\*\* de quelques points de pourcentage (\*1%\*). Ci-dessous est démontré comment cela est fait dans Unsloth : {% columns %} {% column %} \*\*PAS\*\* s'entraîner seulement sur les complétions : \*\*UTILISATEUR :\*\* Bonjour, combien font 2+2 ?\\ \*\*ASSISTANT :\*\* La réponse est 4.\\ \*\*UTILISATEUR :\*\* Bonjour, combien font 3+3 ?\\ \*\*ASSISTANT :\*\* La réponse est 6. {% endcolumn %} {% column %} \*\*Entraînement\*\* sur les complétions seulement : \*\*UTILISATEUR :\*\* ~~Bonjour, combien font 2+2 ?~~\\ \*\*ASSISTANT :\*\* La réponse est 4.\\ \*\*UTILISATEUR :\*\* ~~Bonjour, combien font 3+3 ?~~\\ \*\*ASSISTANT :\*\* La réponse est 6\*\*.\*\* {% endcolumn %} {% endcolumns %} L'article QLoRA indique que \*\*s'entraîner uniquement sur les complétions\*\* augmente considérablement la précision, surtout pour les fine-tunings conversationnels multi-tours ! Nous faisons cela dans nos \[notebooks conversationnels ici\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb). ![](https://unsloth.ai/files/8d35dddcf9a8fd5fc9e41fbde915ec2205640ebf) Pour activer \*\*l'entraînement sur les complétions\*\* dans Unsloth, vous devrez définir les parties instruction et assistant. :sloth: \*Nous prévoyons d'automatiser cela davantage pour vous à l'avenir !\* Pour Llama 3, 3.1, 3.2, 3.3 et les modèles 4, vous définissez les parties comme suit : \`\`\`python from unsloth.chat\_templates import train\_on\_responses\_only trainer = train\_on\_responses\_only( trainer, instruction\_part = "<|start\_header\_id|>user<|end\_header\_id|>\\n\\n", response\_part = "<|start\_header\_id|>assistant<|end\_header\_id|>\\n\\n", ) \`\`\` Pour les modèles Gemma 2, 3, 3n, vous définissez les parties comme suit : \`\`\`python from unsloth.chat\_templates import train\_on\_responses\_only trainer = train\_on\_responses\_only( trainer, instruction\_part = "user\\n", response\_part = "model\\n", ) \`\`\` ## :mag\\\_right:Entraînement uniquement sur les réponses de l'assistant pour les modèles de vision, VLMs Pour les modèles de langage, nous pouvons utiliser \`from unsloth.chat\_templates import train\_on\_responses\_only\` comme décrit précédemment. Pour les modèles de vision, utilisez les arguments supplémentaires dans \`UnslothVisionDataCollator\` comme avant ! {% code overflow="wrap" %} \`\`\`python class UnslothVisionDataCollator : def \_\_init\_\_( self, ... # from unsloth.chat\_templates import train\_on\_responses\_only # trainer = train\_on\_responses\_only( # trainer, # instruction\_part = "<|start\_header\_id|>user<|end\_header\_id|>\\n\\n", # response\_part = "<|start\_header\_id|>assistant<|end\_header\_id|>\\n\\n", # ) train\_on\_responses\_only = False, # ÉQUIVALENT à train\_on\_responses\_only pour les LLM instruction\_part = None, # ÉQUIVALENT à train\_on\_responses\_only(instruction\_part = ...) response\_part = None, # ÉQUIVALENT à train\_on\_responses\_only(response\_part = ...) force\_match = True, # Faire correspondre aussi les retours à la ligne ! ) \`\`\` {% endcode %} Par exemple pour Llama 3.2 Vision : \`\`\`python UnslothVisionDataCollator( model, tokenizer, ... train\_on\_responses\_only = True, instruction\_part = "<|start\_header\_id|>user<|end\_header\_id|>\\n\\n", response\_part = "<|start\_header\_id|>assistant<|end\_header\_id|>\\n\\n", ... ) \`\`\` ## :key: \*\*Éviter le surapprentissage & le sous-apprentissage\*\* ### \*\*Surapprentissage\*\* (Mauvaise généralisation/Trop spécialisé) Le modèle mémorise les données d'entraînement, y compris le bruit statistique, et par conséquent ne parvient pas à généraliser aux données non vues. {% hint style="success" %} Si votre perte d'entraînement descend en dessous de 0.2, votre modèle est probablement \*\*le surapprentissage\*\* — ce qui signifie qu'il peut mal performer sur des tâches non vues. Une astuce simple est la mise à l'échelle de l'alpha LoRA — multipliez simplement la valeur alpha de chaque matrice LoRA par 0.5. Cela réduit effectivement l'impact du fine-tuning. \*\*Ceci est étroitement lié à la fusion / moyenne des poids.\*\*\\ Vous pouvez prendre le modèle de base original (ou instruct), ajouter les poids LoRA, puis diviser le résultat par 2. Cela vous donne un modèle moyenné — ce qui est fonctionnellement équivalent à réduire \`l'alpha\` de moitié. {% endhint %} \*\*Solution :\*\* \* \*\*Ajustez le taux d'apprentissage :\*\* Un taux d'apprentissage élevé conduit souvent au surapprentissage, surtout lors de courts entraînements. Pour des entraînements plus longs, un taux plus élevé peut mieux fonctionner. Il est préférable d'expérimenter les deux pour voir lequel fonctionne le mieux. \* \*\*Réduire le nombre d'époques d'entraînement\*\*. Arrêtez l'entraînement après 1, 2 ou 3 époques. \* \*\*Augmenter\*\* \`weight\_decay\`. Une valeur de \`0.01\` ou \`0.1\` est un bon point de départ. \* \*\*Augmenter\*\* \`lora\_dropout\`. Utilisez une valeur comme \`0.1\` pour ajouter de la régularisation. \* \*\*Augmenter la taille de lot ou les étapes d'accumulation de gradient\*\*. \* \*\*Expansion du jeu de données\*\* - augmentez la taille de votre jeu de données en combinant ou en concaténant des jeux de données open source avec votre jeu de données. Choisissez ceux de meilleure qualité. \* \*\*Arrêt précoce basé sur l'évaluation\*\* - activez l'évaluation et arrêtez lorsque la perte d'évaluation augmente pendant quelques étapes. \* \*\*Mise à l'échelle alpha LoRA\*\* - réduisez l'alpha après l'entraînement et pendant l'inférence - cela rendra le fine-tuning moins prononcé. \* \*\*Moyennage des poids\*\* - ajoutez littéralement le modèle instruct original et le fine-tune puis divisez les poids par 2. ### \*\*Sous-apprentissage\*\* (Trop générique) Le modèle ne parvient pas à capturer les motifs sous-jacents des données d'entraînement, souvent en raison d'une complexité insuffisante ou d'une durée d'entraînement trop courte. \*\*Solution :\*\* \* \*\*Ajustez le taux d'apprentissage :\*\* Si le taux actuel est trop bas, l'augmenter peut accélérer la convergence, surtout pour les courts entraînements. Pour des runs plus longs, essayez plutôt d'abaisser le taux d'apprentissage. Testez les deux approches pour voir laquelle fonctionne le mieux. \* \*\*Augmenter les époques d'entraînement :\*\* Entraînez plus d'époques, mais surveillez la perte de validation pour éviter le surapprentissage. \* \*\*Augmenter le rang LoRA\*\* (\`r\`) et l'alpha : le rang doit au moins être égal au nombre alpha, et le rang doit être plus grand pour les modèles plus petits/jeux de données plus complexes ; il se situe généralement entre 4 et 64. \* \*\*Utiliser un jeu de données plus pertinent pour le domaine\*\* : Assurez-vous que les données d'entraînement sont de haute qualité et directement pertinentes pour la tâche cible. \* \*\*Diminuer la taille de lot à 1\*\*. Cela fera que le modèle se mette à jour plus vigoureusement. {% hint style="success" %} Le fine-tuning n'a pas d'approche "meilleure" unique, seulement des meilleures pratiques. L'expérimentation est la clé pour trouver ce qui fonctionne pour vos besoins spécifiques. Nos notebooks définissent automatiquement des paramètres optimaux basés sur de nombreuses recherches et nos expériences, vous offrant un excellent point de départ. Bon fine-tuning ! {% endhint %} \*\*\*Remerciements :\*\* Un énorme merci à\* \[\*Eyera\*\](https://huggingface.co/Orenguteng) \*pour avoir contribué à ce guide !\* --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md). # Vision Reinforcement Learning (VLM RL) Unsloth now supports vision/multimodal RL with \[Qwen3-VL\](/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune.md), \[Gemma 3\](/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md) and more. Due to Unsloth's unique \[weight sharing\](/docs/get-started/reinforcement-learning-rl-guide.md#what-unsloth-offers-for-rl) and custom kernels, Unsloth makes VLM RL \*\*1.5–2× faster,\*\* uses \*\*90% less VRAM\*\*, and enables \*\*15× longer context\*\* lengths than FA2 setups, with no accuracy loss. This update also introduces Qwen's \[GSPO\](#gspo-rl) algorithm. Unsloth can train Qwen3-VL-8B with GSPO/GRPO on a free Colab T4 GPU. Other VLMs work too, but may need larger GPUs. Gemma requires newer GPUs than T4 because vLLM \[restricts to Bfloat16\](/docs/models/tutorials/gemma-3-how-to-run-and-fine-tune.md#unsloth-fine-tuning-fixes), thus we recommend NVIDIA L4 on Colab. Our notebooks solve numerical math problems involving images and diagrams: \* \*\*Qwen-3 VL-8B\*\* (vLLM inference)\*\*:\*\* \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) \* \*\*Qwen-2.5 VL-7B\*\* (vLLM inference)\*\*:\*\* \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2\_5\_7B\_VL\_GRPO.ipynb) •\[ Kaggle\](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen2\_5\_7B\_VL\_GRPO.ipynb\\&accelerator=nvidiaTeslaT4) \* \*\*Gemma-3-4B\*\* (Unsloth inference): \[Colab\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\)-Vision-GRPO.ipynb) We have also added vLLM VLM integration into Unsloth natively, so all you have to do to use vLLM inference is enable the \`fast\_inference=True\` flag when initializing the model. Special thanks to \[Sinoué GAD\](https://github.com/unslothai/unsloth/pull/2752) for providing the \[first notebook\](https://github.com/GAD-cell/vlm-grpo/blob/main/examples/VLM\_GRPO\_basic\_example.ipynb) that made integrating VLM RL easier! This VLM support also integrates our latest update for even more memory efficient + faster RL including our \[Standby feature\](/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl.md#unsloth-standby), which uniquely limits speed degradation compared to other implementations. {% hint style="info" %} You can only use \`fast\_inference\` for VLMs supported by vLLM. Some models, like Llama 3.2 Vision thus only can run without vLLM, but they still work in Unsloth. {% endhint %} \`\`\`python os.environ\['UNSLOTH\_VLLM\_STANDBY'\] = '1' # To enable memory efficient GRPO with vLLM model, tokenizer = FastVisionModel.from\_pretrained( model\_name = "Qwen/Qwen2.5-VL-7B-Instruct", max\_seq\_length = 16384, #Must be this large to fit image in context load\_in\_4bit = True, # False for LoRA 16bit fast\_inference = True, # Enable vLLM fast inference gpu\_memory\_utilization = 0.8, # Reduce if out of memory ) \`\`\` It is also important to note, that vLLM does not support LoRA for vision/encoder layers, thus set \`finetune\_vision\_layers = False\` when loading a LoRA adapter.\\ However you CAN train the vision layers as well if you use inference via transformers/Unsloth. \`\`\`python # Add LoRA adapter to the model for parameter efficient fine tuning model = FastVisionModel.get\_peft\_model( model, finetune\_vision\_layers = False,# fast\_inference doesn't support finetune\_vision\_layers yet :( finetune\_language\_layers = True, # False if not finetuning language layers finetune\_attention\_modules = True, # False if not finetuning attention layers finetune\_mlp\_modules = True, # False if not finetuning MLP layers r = lora\_rank, # Choose any number > 0 ! Suggested 8, 16, 32, 64, 128 lora\_alpha = lora\_rank\*2, # \*2 speeds up training use\_gradient\_checkpointing = "unsloth", # Reduces memory usage random\_state = 3407, ) \`\`\` ## :butterfly:Qwen 2.5 VL Vision RL Issues and Quirks During RL for Qwen 2.5 VL, you might see the following inference output: {% code overflow="wrap" %} \`\`\` addCriterion \\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n\\n addCriterion\\n\\n 自动生成\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n addCriterion\\n\\n\\n addCriterion\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n\\n \`\`\` {% endcode %} This was \[reported\](https://github.com/QwenLM/Qwen2.5-VL/issues/759) as well in Qwen2.5-VL-7B-Instruct output unexpected results "addCriterion". In fact we see this as well! We tried both non Unsloth, bfloat16 and float16 machines and other things, but it appears still. For example item 165 ie \`train\_dataset\[165\]\` from the \[AI4Math/MathVista\](https://huggingface.co/datasets/AI4Math/MathVista) dataset is below: {% code overflow="wrap" %} \`\`\` Figure is an overhead view of the path taken by a race car driver as his car collides with the racetrack wall. Just before the collision, he is traveling at speed $v\_i=70 \\mathrm{~m} / \\mathrm{s}$ along a straight line at $30^{\\circ}$ from the wall. Just after the collision, he is traveling at speed $v\_f=50 \\mathrm{~m} / \\mathrm{s}$ along a straight line at $10^{\\circ}$ from the wall. His mass $m$ is $80 \\mathrm{~kg}$. The collision lasts for $14 \\mathrm{~ms}$. What is the magnitude of the average force on the driver during the collision? \`\`\` {% endcode %} ![](https://unsloth.ai/files/uaJpqsjtTh4s6AHEaHZQ) And then we get the above gibberish output. One could add a reward function to penalize the addition of addCriterion, or penalize gibberish outputs. However, the other approach is to train it for longer. For example only after 60 steps ish do we see the model actually learning via RL: ![](https://unsloth.ai/files/VaNkR5oCVwVstrhPFnvj) {% hint style="success" %} Forcing \`<|assistant|>\` during generation will reduce the occurrences of these gibberish results as expected since this is an Instruct model, however it's still best to add a reward function to penalize bad generations, as described in the next section. {% endhint %} ## :medal:Reward Functions to reduce gibberish To penalize \`addCriterion\` and gibberish outputs, we edited the reward function to penalize too much of \`addCriterion\` and newlines. \`\`\`python def formatting\_reward\_func(completions,\*\*kwargs): import re thinking\_pattern = f'{REASONING\_START}(.\*?){REASONING\_END}' answer\_pattern = f'{SOLUTION\_START}(.\*?){SOLUTION\_END}' scores = \[\] for completion in completions: score = 0 thinking\_matches = re.findall(thinking\_pattern, completion, re.DOTALL) answer\_matches = re.findall(answer\_pattern, completion, re.DOTALL) if len(thinking\_matches) == 1: score += 1.0 if len(answer\_matches) == 1: score += 1.0 # Fix up addCriterion issues # See https://docs.unsloth.ai/new/vision-reinforcement-learning-vlm-rl#qwen-2.5-vl-vision-rl-issues-and-quirks # Penalize on excessive addCriterion and newlines if len(completion) != 0: removal = completion.replace("addCriterion", "").replace("\\n", "") if (len(completion)-len(removal))/len(completion) >= 0.5: score -= 2.0 scores.append(score) return scores \`\`\` ## :checkered\\\_flag:GSPO Reinforcement Learning This update in addition adds GSPO (\[Group Sequence Policy Optimization\](https://arxiv.org/abs/2507.18071)) which is a variant of GRPO made by the Qwen team at Alibaba. They noticed that GRPO implicitly results in importance weights for each token, even though explicitly advantages do not scale or change with each token. This lead to the creation of GSPO, which now assigns the importance on the sequence likelihood rather than the individual token likelihoods of the tokens. The difference between these two algorithms can be seen below, both from the GSPO paper from Qwen and Alibaba: ![](https://unsloth.ai/files/7ntPBNOCiL5uh506pLzt) GRPO Algorithm, Source: [Qwen](https://arxiv.org/abs/2507.18071) ![](https://unsloth.ai/files/eyxi56SIssHb6e4xuA7L) GSPO algorithm, Source: [Qwen](https://arxiv.org/abs/2507.18071) In Equation 1, it can be seen that the advantages scale each of the rows into the token logprobs before that tensor is sumed. Essentially, each token is given the same scaling even though that scaling was given to the entire sequence rather than each individual token. A simple diagram of this can be seen below: ![](https://unsloth.ai/files/ebkGBYfegbkcvTt89vou) GRPO Logprob Ratio row wise scaled with advantages Equation 2 shows that the logprob ratios for each sequence is summed and exponentiated after the Logprob ratios are computed, and only the resulting now sequence ratios get row wise multiplied by the advantages. ![](https://unsloth.ai/files/jPCxEjDztfNea4N75aF3) GSPO Sequence Ratio row wise scaled with advantages Enabling GSPO is simple, all you need to do is set the \`importance\_sampling\_level = "sequence"\` flag in the GRPO config. \`\`\`python training\_args = GRPOConfig( output\_dir = "vlm-grpo-unsloth", per\_device\_train\_batch\_size = 8, gradient\_accumulation\_steps = 4, learning\_rate = 5e-6, adam\_beta1 = 0.9, adam\_beta2 = 0.99, weight\_decay = 0.1, warmup\_ratio = 0.1, lr\_scheduler\_type = "cosine", optim = "adamw\_8bit", # beta = 0.00, epsilon = 3e-4, epsilon\_high = 4e-4, num\_generations = 8, max\_prompt\_length = 1024, max\_completion\_length = 1024, log\_completions = False, max\_grad\_norm = 0.1, temperature = 0.9, # report\_to = "none", # Set to "wandb" if you want to log to Weights & Biases num\_train\_epochs = 2, # For a quick test run, increase for full training report\_to = "none" # GSPO is below: importance\_sampling\_level = "sequence", # Dr GRPO / GAPO etc loss\_type = "dr\_grpo", ) \`\`\` Overall, Unsloth now with VLM vLLM fast inference enables for both 90% reduced memory usage but also 1.5-2x faster speed with GRPO and GSPO! If you'd like to read more about reinforcement learning, check out out RL guide: \[Reinforcement Learning\](/docs/get-started/reinforcement-learning-rl-guide.md) \*\*\*Authors:\*\* A huge thank you to\* \[\*Keith\*\](https://www.linkedin.com/in/keith-truongcao-7bb84a23b/) \*and\* \[\*Datta\*\](https://www.linkedin.com/in/datta0/) \*for contributing to this article!\* --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/new/studio/data-recipe.md). # Unsloth Data Recipes Unsloth Studio's Data Recipes lets you upload documents like PDFs or CSVs files and transforms them into useable / synthetic datasets. Create and edit datasets visually via a graph-node workflow. This guide will get you started with the basics before you dive into Unsloth Data Recipes. ![](https://unsloth.ai/files/WLEs48rQ0pkULMPgGcwU) \### How Data Recipes works Data Recipes follows the same basic path. You open the recipes page, create or pick a recipe, build the workflow in the editor, validate it run a preview, then run the full dataset once the output looks right. Add seed data and generation blocks, validate the workflow, preview sample output, then run a full dataset build. Unsloth Data Recipes is powered by \*\*NVIDIA Nemo\*\* \[\*\*Data Designer\*\*\](https://github.com/NVIDIA-NeMo/DataDesigner). ![](https://unsloth.ai/files/9h99m1UMA8mx5wJ4rcxB) Example of generating dataset and fine-tuning a model At a glance a usual workflow should look like this: 1. Open the recipes page. 2. Create a new recipe or open an existing one. 3. Add blocks to define your dataset workflow. 4. Click \*\*Validate\*\* to catch configuration issues early. 5. Run a preview to inspect sample rows quickly. 6. Run a full dataset build when the recipe is ready. 7. Review progress and output live in graph or in \*\*Executions\*\* view for mode details. 8. Select the resulting dataset in \*\*Unsloth\*\* and fine tune a model. ### Get Started The recipes page is the main entry point. Recipes are stored locally in the browser, so you come back to saved work later. From here, you can create a blank recipe or open a guided learning recipe. {% hint style="info" %} Recipes can be exported and imported, so it is easy to share workflows with other Unsloth users :tada:. If you are trying to build a specific dataset pattern, ask in Unsloth Discord. Someone may already have a recipe they can share. {% endhint %} ![](https://unsloth.ai/files/80CkTFi1G0KP6n0Vftod) Recipes landing page If you are new to concept of workflows, learning recipes are the fastest way to see how seed data, prompts, expressions, and validators fit together in one working example. If you already know the shape of dataset you want, starting empty is usually quicker. #### Choose a starting path | If you want to: | Start with: | | | --- | --- | --- | | **Build a custom workflow quickly** | **Start Empty** | | | **Learn the product from an example** | **Start from Learning Recipe** | | | **Continue previous work** | **Open a saved recipe** | | \### What you build in the editor The editor is where the recipe takes shape. You add blocks from the block sheet, configure them in dialogs, connect them on the canvas, and then validate or run the workflow. ![](https://unsloth.ai/files/LDOek8qFkSkUvUpqReVG) Example of building product description workflow {% columns %} {% column %} The editor has a few core parts: \* The recipe header, where you rename the recipe and switch between \*\*Editor\*\* and \*\*Executions\*\* \* The canvas, where the recipe graph is shown \* The block sheet, where you add new blocks \* Configuration dialogs, where you define prompts, references, model aliases, validators and seed settings. \* The floating \*\*Run\*\* and \*\*Validate\*\* controls \* need to add more here {% endcolumn %} {% column %} The most common blocks in reciper are: \* \*\*Seed\*\* for input data from hugginface, local structured files (or unstructured documents that get chunked into rows. \* \*\*LLM + Models\*\* for providers, model configs, LLM generation blocks, and shared tool profiles. \* \*\*Expression\*\* for jinja2-based transforms that do not require an LLM call. \* \*\*Validators\*\* for filtering bad generated code with built in linters for Python, SQL, and Javascript/Typescript. \* \*\*Samplers\*\* for deterministic columns such as categories and subcategories. {% endcolumn %} {% endcolumns %} ### How references work Most blocks that produce data (with some exceptions) becomes a reference for later blocks. That is one of the main ideas behind Data Recipes. You create a value once, then reuse it in prompts, expressions, structured outputs, and validation steps. {% hint style="info" %} Jinja Expressions help you work with values that arleady exist in the recipe. You can reference nested fields like \`{{customer.first\_name}}\` , join values like \`{{customer.first\_name}} {{customer.last\_name}}\` and add conditional logic with patterns such as \`{% if condition %}...{% endif %}\` {% endhint %} ![](https://unsloth.ai/files/zXnXJ4Jd6byRbSJYrvCd) Example of references shown in the editor For example: \* A category block named \`domain\` can be references as \`{{ domain }}\` \* a seed column can be used directly in an LLM prompt, the columns in your seed data (eg. HF dataset columns, csv) \* a structured LLM output can expose fileds for later prompts \* an expression block can combine earlyier values without another model call ### What happens after? Preview runs are for quick iteration. They return sample rows and analysis in the editor so you can inspect the generated data before commiting to a full run. Full runs create a persisted local dataset artifact. That output later appears in Unsloth's local dataset picker, where you can inspect it again and use it for fine-tuning. Optionally you can publish your dataset to you hugginface repo. ### Core building blocks {% columns %} {% column %} ![](https://unsloth.ai/files/gLlIppqXVaQ5QuPsxN62) Core building blocks {% endcolumn %} {% column %} ![](https://unsloth.ai/files/ltMnq2jkGXm6jtkkiifY) Model and LLM blocks {% endcolumn %} {% endcolumns %} #### Model setup is split into two usable layers: \* \*\*Model provider\*\* defines the endpoint and authentifcation \* \*\*Model Config\*\* defines the model name and inference settings This setup works with hosted providers, self-hosted endpoints, \`vLLM\` , \`llama.cpp\` , or any OpenAI-compatible API that you run outside Unsloth. {% hint style="info" %} Recipes are not limited to one model. You can add multiple \*\*Model providers\*\* and \*\*Model config\*\* blocks, then use different models for different steps, such as one for coding and another for general text tasks. {% endhint %} After model setup, you can use Four LLM block types: | Block | Output | Best for | | -------------- | ----------------- | ----------------------------------------------------------- | | LLM Text | Free-form text | Instructions, explanations, conversations, and descriptions | | LLM Structured | JSON | Output that need fixed fields and predictable structure | | LLM Code | Code | Python, SQL, Typescript and other code generation tasks | | LLM Judge | Scored evaluation | Grading outputs with one or more user-defined score | #### Tool Profiles {% columns %} {% column %} Tool profile blocks defines shared MCP based tool access for one or more LLM blocks. Use them when a generation step needs tools, such as looking up code documentation through \`Context7\`. Image to the left shows Context7 MCP added and configured in Tool Profile block dialog: {% endcolumn %} {% column %} ![](https://unsloth.ai/files/L2nK8KRo7A0U3ygadHfF) {% endcolumn %} {% endcolumns %} #### Validators {% columns %} {% column %} Validor block primarly target LLM code block by running generated code outputs through Linter and syntax validation, this helps you keep bad or invalid code rows out of the final dataset by filtering them out. The built-in options cover Python, SQL, and JavaScript/TypeScript validation. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/hQcs8p9q9MC5gjxgDc4A) {% endcolumn %} {% endcolumns %} ### Validate, preview and run Once the recipe workflow is in place, the next step is execution. The reccomended pattern is: validate first, preview for quick feedback and inspect the generated data in executions view, then run the full dataset when you feel the output satisfies your plan. Use the execution controls in third order: {% stepper %} {% step %} #### Validate Click \*\*Validate\*\* to catch configuration issues. {% endstep %} {% step %} #### Preview Run a preview to inspect sample rows and analysis {% endstep %} {% step %} #### Refine Refine prompts, references, seed settings, or validators. Iterate untill you feel satisfied with generated data {% endstep %} {% step %} #### Run the full dataset build {% endstep %} {% endstepper %} ![](https://unsloth.ai/files/HYTosiOXnIW7MBaHHy88) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/new/studio/data-recipe.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/datasets-guide.md). # Guide des jeux de données ## Qu'est-ce qu'un jeu de données ? Pour les LLM, les jeux de données sont des collections de données qui peuvent être utilisées pour entraîner nos modèles. Afin d'être utiles à l'entraînement, les données textuelles doivent être dans un format pouvant être tokenisé. Vous apprendrez aussi comment \[utiliser des jeux de données dans Unsloth\](#applying-chat-templates-with-unsloth). L'une des parties clés de la création d'un jeu de données est votre \[modèle de chat\](/docs/fr/notions-de-base/chat-templates.md) et la manière dont vous allez le concevoir. La tokenisation est également importante car elle découpe le texte en jetons, qui peuvent être des mots, des sous-mots ou des caractères, afin que les LLM puissent le traiter efficacement. Ces jetons sont ensuite convertis en embeddings et ajustés pour aider le modèle à comprendre le sens et le contexte. ### Format des données Pour permettre le processus de tokenisation, les jeux de données doivent être dans un format lisible par un tokenizer. | Format | Description | Type d'entraînement | | --- | --- | --- | | Corpus brut | Texte brut provenant d'une source telle qu'un site web, un livre ou un article. | Préentraînement continu (CPT) | | Consigne | Instructions à suivre pour le modèle et un exemple de la sortie visée. | Ajustement fin supervisé (SFT) | | Conversation | Conversation à plusieurs tours entre un utilisateur et un assistant IA. | Ajustement fin supervisé (SFT) | | RLHF | Conversation entre un utilisateur et un assistant IA, les réponses de l'assistant étant classées par un script, un autre modèle ou un évaluateur humain. | Apprentissage par renforcement (RL) | {% hint style="info" %} Il convient de noter qu'il existe différents styles de format pour chacun de ces types. {% endhint %} ## Pour commencer Avant de mettre nos données en forme, nous voulons identifier les éléments suivants : {% stepper %} {% step %} Objectif du jeu de données Connaître l'objectif du jeu de données nous aidera à déterminer quelles données nous devons utiliser et quel format adopter. L'objectif peut être d'adapter un modèle à une nouvelle tâche, comme le résumé, ou d'améliorer la capacité d'un modèle à jouer le rôle d'un personnage spécifique. Par exemple : \* Dialogues basés sur le chat (Q\\&R, apprentissage d'une nouvelle langue, support client, conversations). \* Tâches structurées (\[classification\](https://colab.research.google.com/github/timothelaborie/text\_classification\_scripts/blob/main/unsloth\_classification.ipynb), résumé, tâches de génération). \* Données spécifiques à un domaine (médical, finance, technique). {% endstep %} {% step %} Style de sortie Le style de sortie nous indiquera quelles sources de données nous utiliserons pour atteindre la sortie souhaitée. Par exemple, le type de sortie que vous souhaitez obtenir peut être du JSON, du HTML, du texte ou du code. Ou peut-être souhaitez-vous qu'il soit en espagnol, en anglais, en allemand, etc. {% endstep %} {% step %} Source des données Lorsque nous connaissons l'objectif et le style des données dont nous avons besoin, nous devons analyser la qualité et la \[quantité\](#how-big-should-my-dataset-be) des données. Hugging Face et Wikipédia sont d'excellentes sources de jeux de données et Wikipédia est particulièrement utile si vous cherchez à entraîner un modèle à apprendre une langue. La source des données peut être un fichier CSV, un PDF ou même un site web. Vous pouvez également \[générer de manière synthétique\](#synthetic-data-generation) des données, mais un soin particulier est nécessaire pour s'assurer que chaque exemple est de haute qualité et pertinent. {% endstep %} {% endstepper %} {% hint style="success" %} L'une des meilleures façons de créer un meilleur jeu de données consiste à le combiner avec un jeu de données plus général issu de Hugging Face, comme ShareGPT, afin de rendre votre modèle plus intelligent et plus diversifié. Vous pourriez également ajouter \[des données générées synthétiquement\](#synthetic-data-generation). {% endhint %} ## 🦥 Recettes de données Unsloth \[Recettes de données Unsloth\](/docs/fr/nouveau/studio/data-recipe.md) vous permet de téléverser des documents comme des PDF ou des fichiers CSV et de les transformer en jeux de données utilisables. Créez et modifiez visuellement des jeux de données via un workflow de graphe de nœuds. La page des recettes est le point d’entrée principal. Les recettes sont stockées localement dans le navigateur, ce qui vous permet de retrouver plus tard votre travail enregistré. À partir de là, vous pouvez créer une recette vierge ou ouvrir une recette d’apprentissage guidée. ![](https://unsloth.ai/files/180bd7cd3cbd0cc95559d28cb19d6385f8c94f50) Data Recipes suit le même chemin de base. Vous ouvrez la page des recettes, créez ou choisissez une recette, construisez le flux de travail dans l'éditeur, le validez, exécutez un aperçu, puis lancez le jeu de données complet une fois que la sortie semble correcte. Ajoutez des données de départ et des blocs de génération, validez le flux de travail, prévisualisez un exemple de sortie, puis lancez la création complète du jeu de données. À première vue, un flux de travail habituel devrait ressembler à ceci : 1. Ouvrez la page des recettes. 2. Créez une nouvelle recette ou ouvrez-en une existante. 3. Ajoutez des blocs pour définir le flux de travail de votre jeu de données. 4. Cliquez sur \*\*Validez\*\* pour détecter rapidement les problèmes de configuration. 5. Lancez un aperçu pour examiner rapidement des lignes d’exemple. 6. Lancez une génération complète du jeu de données lorsque la recette est prête. 7. Suivez la progression et la sortie en direct dans le graphe ou dans \*\*Exécutions\*\* vue pour plus de détails. 8. Sélectionnez le jeu de données résultant dans Unsloth et ajustez finement un modèle. En savoir plus : {% content-ref url="/pages/ec921eec2a37820cfd0cf432328d619c9fa1080b" %} \[Data Recipes\](/docs/fr/nouveau/studio/data-recipe.md) {% endcontent-ref %} ## Mise en forme des données Lorsque nous avons identifié les critères pertinents et collecté les données nécessaires, nous pouvons alors mettre nos données en forme dans un format lisible par machine, prêt pour l'entraînement. ### Formats de données courants pour l'entraînement des LLM Pour \[\*\*préentraînement continu\*\*\](/docs/fr/notions-de-base/continued-pretraining.md), nous utilisons un format de texte brut sans structure spécifique : \`\`\`json "text": "Les pâtes carbonara sont un plat traditionnel romain. La sauce est préparée en mélangeant des œufs crus avec du fromage Pecorino Romano râpé et du poivre noir. Les pâtes chaudes sont ensuite mélangées avec du guanciale croustillant (joue de porc salée) et le mélange d'œufs, créant une sauce crémeuse grâce à la chaleur résiduelle. Malgré l'idée reçue, la vraie carbonara ne contient jamais de crème ni d'ail. Le plat est probablement originaire de Rome au milieu du XXe siècle, bien que ses origines exactes soient débattues..." \`\`\` Ce format préserve le flux naturel du langage et permet au modèle d'apprendre à partir d'un texte continu. Si nous adaptons un modèle à une nouvelle tâche, et que nous voulons que le modèle génère du texte en un seul tour en se basant sur un ensemble précis d'instructions, nous pouvons utiliser \*\*Instruction\*\* format en \[style Alpaca\](https://docs.unsloth.ai/basics/tutorial-how-to-finetune-llama-3-and-use-in-ollama#id-6.-alpaca-dataset) \`\`\`json "Instruction": "Tâche que nous voulons que le modèle réalise." "Input": "Optionnel, mais utile ; il s'agira essentiellement de la requête de l'utilisateur." "Output": "Le résultat attendu de la tâche et la sortie du modèle." \`\`\` Lorsque nous voulons plusieurs tours de conversation, nous pouvons utiliser le format ShareGPT : \`\`\`json { "conversations": \[ { "from": "human", "value": "Pouvez-vous m'aider à faire des pâtes carbonara ?" }, { "from": "gpt", "value": "Souhaitez-vous la recette romaine traditionnelle ou une version plus simple ?" }, { "from": "human", "value": "La version traditionnelle, s'il vous plaît" }, { "from": "gpt", "value": "La vraie carbonara romaine utilise seulement quelques ingrédients : pâtes, guanciale, œufs, Pecorino Romano et poivre noir. Souhaitez-vous la recette détaillée ?" } \] } \`\`\` Le format modèle utilise les clés d'attributs « from »/« value » et les messages alternent entre \`humain\`et \`gpt\`, ce qui permet un flux naturel du dialogue. L'autre format courant est le format ChatML d'OpenAI, et c'est celui que Hugging Face utilise par défaut. C'est probablement le format le plus utilisé, et il alterne entre \`user\` et \`assistant\` \`\`\` { "messages": \[ { "role": "user", "content": "Combien font 1+1 ?" }, { "role": "assistant", "content": "C'est 2 !" }, \] } \`\`\` ### Application des modèles de chat avec Unsloth Pour les jeux de données qui suivent généralement le format chatml courant, le processus de préparation du jeu de données pour l'entraînement ou le fine-tuning consiste en quatre étapes simples : \* Consultez les modèles de chat actuellement pris en charge par Unsloth :\\\\ \`\`\` from unsloth.chat\_templates import CHAT\_TEMPLATES print(list(CHAT\_TEMPLATES.keys())) \`\`\` \\ Cela affichera la liste des modèles actuellement pris en charge par Unsloth. Voici un exemple de sortie :\\\\ \`\`\` \['unsloth', 'zephyr', 'chatml', 'mistral', 'llama', 'vicuna', 'vicuna\_old', 'vicuna old', 'alpaca', 'gemma', 'gemma\_chatml', 'gemma2', 'gemma2\_chatml', 'llama-3', 'llama3', 'phi-3', 'phi-35', 'phi-3.5', 'llama-3.1', 'llama-31', 'llama-3.2', 'llama-3.3', 'llama-32', 'llama-33', 'qwen-2.5', 'qwen-25', 'qwen25', 'qwen2.5', 'phi-4', 'gemma-3', 'gemma3'\] \`\`\` \\\\ \* Utilisez \`get\_chat\_template\` pour appliquer le bon modèle de chat à votre tokenizer :\\\\ \`\`\` from unsloth.chat\_templates import get\_chat\_template tokenizer = get\_chat\_template( tokenizer, chat\_template = "gemma-3", # remplacez ceci par le bon nom de chat\_template ) \`\`\` \\\\ \* Définissez votre fonction de mise en forme. Voici un exemple :\\\\ \`\`\` def formatting\_prompts\_func(examples): convos = examples\["conversations"\] texts = \[tokenizer.apply\_chat\_template(convo, tokenize = False, add\_generation\_prompt = False) for convo in convos\] return { "text" : texts, } \`\`\` \\ \\ Cette fonction parcourt votre jeu de données en appliquant le modèle de chat que vous avez défini à chaque exemple.\\\\ \* Enfin, chargeons le jeu de données et appliquons les modifications requises à notre jeu de données : \\\\ \`\`\` # Importer et charger le jeu de données from datasets import load\_dataset dataset = load\_dataset("repo\_name/dataset\_name", split = "train") # Appliquez la fonction de mise en forme à votre jeu de données à l'aide de la méthode map dataset = dataset.map(formatting\_prompts\_func, batched = True,) \`\`\` \\ Si votre jeu de données utilise le format ShareGPT avec les clés « from »/« value » au lieu du format ChatML « role »/« content », vous pouvez utiliser la \`standardize\_sharegpt\` fonction pour le convertir d'abord. Le code révisé ressemblera désormais à ceci :\\ \\\\ \`\`\` # Importer le jeu de données from datasets import load\_dataset dataset = load\_dataset("mlabonne/FineTome-100k", split = "train") # Convertissez votre jeu de données au format « role »/« content » si nécessaire from unsloth.chat\_templates import standardize\_sharegpt dataset = standardize\_sharegpt(dataset) # Appliquez la fonction de mise en forme à votre jeu de données à l'aide de la méthode map dataset = dataset.map(formatting\_prompts\_func, batched = True,) \`\`\` ### Mise en forme des données Q\\&R \*\*Q :\*\* Comment puis-je utiliser le format d'instruction Alpaca ? \*\*R :\*\* Si votre jeu de données est déjà formaté au format Alpaca, suivez alors les étapes de mise en forme comme indiqué dans le notebook Llama3.1 \[notebook \](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1\_\\(8B\\)-Alpaca.ipynb#scrollTo=LjY75GoYUCB8). Si vous devez convertir vos données au format Alpaca, une approche consiste à créer un script Python pour traiter vos données brutes. Si vous travaillez sur une tâche de résumé, vous pouvez utiliser un LLM local pour générer des instructions et des sorties pour chaque exemple. \*\*Q :\*\* Dois-je toujours utiliser la méthode standardize\\\_sharegpt ? \*\*R :\*\* N'utilisez la méthode standardize\\\_sharegpt que si votre jeu de données cible est au format sharegpt, mais que votre modèle attend à la place un format ChatML. \\ \*\*Q :\*\* Pourquoi ne pas utiliser la fonction apply\\\_chat\\\_template fournie avec le tokenizer. \*\*R :\*\* Le \`chat\_template\` de l'attribut lorsqu'un modèle est d'abord téléversé par les propriétaires initiaux du modèle contient parfois des erreurs et peut prendre du temps à être mis à jour. En revanche, chez Unsloth, nous vérifions et corrigeons minutieusement toute erreur dans le \`chat\_template\` pour chaque modèle lorsque nous téléversons les versions quantifiées dans nos dépôts. De plus, nos \`get\_chat\_template\` et \`apply\_chat\_template\` méthodes offrent des fonctionnalités avancées de manipulation des données, qui sont entièrement documentées sur notre documentation des modèles de chat \[page\](https://docs.unsloth.ai/basics/chat-templates). \*\*Q :\*\* Et si mon modèle n'est actuellement pas pris en charge par Unsloth ? \*\*R :\*\* Soumettez une demande de fonctionnalité sur le forum des issues GitHub d'Unsloth \[forum\](https://github.com/unslothai/unsloth). Comme solution temporaire, vous pouvez aussi utiliser la fonction apply\\\_chat\\\_template du tokenizer jusqu'à ce que votre demande de fonctionnalité soit approuvée et fusionnée. ## Génération de données synthétiques Vous pouvez également utiliser n'importe quel LLM local comme Llama 3.3 (70B) ou GPT 4.5 d'OpenAI pour générer des données synthétiques. En général, il vaut mieux utiliser un modèle plus grand comme Llama 3.3 (70B) afin de garantir la meilleure qualité de sortie. Vous pouvez utiliser directement des moteurs d'inférence comme vLLM, Ollama ou llama.cpp pour générer des données synthétiques, mais cela demandera un peu de travail manuel pour les collecter et demander davantage de données. Il y a 3 objectifs pour les données synthétiques : \* Produire des données entièrement nouvelles — soit à partir de zéro, soit à partir de votre jeu de données existant \* Diversifier votre jeu de données afin que votre modèle ne \[surapprenne\](/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide.md#avoiding-overfitting-and-underfitting) et ne devienne pas trop spécifique \* Compléter les données existantes, par exemple structurer automatiquement votre jeu de données dans le bon format choisi ### Utiliser Unsloth pour les données synthétiques Vous pouvez facilement téléverser n'importe quelles données non structurées ou structurées dans \[Recettes de données\](/docs/fr/nouveau/studio/data-recipe.md) et il les convertira automatiquement en un jeu de données exploitable / synthétique. Plus de détails dans \[notre guide\](/docs/fr/nouveau/studio/data-recipe.md). ![](https://unsloth.ai/files/cdbdb9ee45f65c5b9d2621ee85677f55f5ff4b08) \### Utiliser un LLM local ou ChatGPT pour les données synthétiques Votre objectif est de demander au modèle de générer et de traiter des données Q\\&R dans le format que vous avez spécifié. Le modèle devra apprendre la structure que vous avez fournie ainsi que le contexte ; veillez donc à disposer d'au moins 10 exemples de données déjà existants. Exemples d'invites : \* \*\*Invite pour générer davantage de dialogues à partir d'un jeu de données existant\*\*: En utilisant l'exemple de jeu de données que j'ai fourni, suivez la structure et générez des conversations basées sur les exemples. \* \*\*Invite si vous n'avez pas de jeu de données\*\*: {% code overflow="wrap" %} \`\`\` Créez 10 exemples d'avis produits pour Coca-Cola classés comme positifs, négatifs ou neutres. \`\`\` {% endcode %} \* \*\*Invite pour un jeu de données sans mise en forme\*\*: {% code overflow="wrap" %} \`\`\` Structurez mon jeu de données pour qu'il soit au format Q&R ChatML pour le fine-tuning. Ensuite, générez 5 exemples de données synthétiques avec le même sujet et le même format. \`\`\` {% endcode %} Il est recommandé de vérifier la qualité des données générées afin de supprimer ou d'améliorer les réponses hors sujet ou de mauvaise qualité. Selon votre jeu de données, il peut également être nécessaire de l'équilibrer sur de nombreux aspects afin que votre modèle ne surapprenne pas. Vous pouvez ensuite réinjecter ce jeu de données nettoyé dans votre LLM pour régénérer des données, cette fois avec encore plus de guidance. ## FAQ + conseils sur les jeux de données ### Quelle taille devrait avoir mon jeu de données ? Nous recommandons généralement d'utiliser au minimum absolu 100 lignes de données pour le fine-tuning afin d'obtenir des résultats raisonnables. Pour des performances optimales, un jeu de données de plus de 1 000 lignes est préférable ; dans ce cas, davantage de données conduit généralement à de meilleurs résultats. Si votre jeu de données est trop petit, vous pouvez également ajouter des données synthétiques ou ajouter un jeu de données provenant de Hugging Face pour le diversifier. Cependant, l'efficacité de votre modèle ajusté dépend fortement de la qualité du jeu de données ; veillez donc à nettoyer et préparer vos données de manière approfondie. ### Comment dois-je structurer mon jeu de données si je veux ajuster finement un modèle de raisonnement ? Si vous voulez ajuster finement un modèle qui possède déjà des capacités de raisonnement, comme les versions distillées de DeepSeek-R1 (par ex. DeepSeek-R1-Distill-Llama-8B), vous devrez tout de même suivre des paires question/tâche et réponse ; cependant, pour votre réponse, vous devrez la modifier afin qu'elle inclue un processus de raisonnement/chaîne de pensée et les étapes suivies pour dériver la réponse.\\ \\ Pour un modèle qui ne possède pas de raisonnement et que vous voulez entraîner afin qu'il acquière ensuite des capacités de raisonnement, vous devrez utiliser un jeu de données standard mais cette fois sans raisonnement dans ses réponses. Ce processus d'entraînement est connu sous le nom de \[Apprentissage par renforcement et GRPO\](/docs/fr/commencer/reinforcement-learning-rl-guide.md). ### Plusieurs jeux de données Si vous avez plusieurs jeux de données pour le fine-tuning, vous pouvez soit : \* Uniformiser le format de tous les jeux de données, les combiner en un seul jeu de données et effectuer le fine-tuning sur ce jeu de données unifié. \* Utilisez le \[Jeux de données multiples\](https://colab.research.google.com/drive/1njCCbE1YVal9xC83hjdo2hiGItpY\_D6t?usp=sharing) notebook pour effectuer directement le fine-tuning sur plusieurs jeux de données. ### Puis-je ajuster finement le même modèle plusieurs fois ? Vous pouvez ajuster finement plusieurs fois un modèle déjà ajusté, mais il est préférable de combiner tous les jeux de données et d'effectuer le fine-tuning en un seul processus à la place. Entraîner un modèle déjà ajusté peut potentiellement modifier la qualité et les connaissances acquises lors du précédent processus de fine-tuning. ## Utilisation des jeux de données dans Unsloth ### Jeu de données Alpaca Voir un exemple d'utilisation du jeu de données Alpaca dans Unsloth sur Google Colab : ![](https://unsloth.ai/files/d92d20b20285d5ca74ab275620c603add74a2ecf) Nous allons maintenant utiliser le jeu de données Alpaca créé en appelant GPT-4 lui-même. Il s'agit d'une liste de 52 000 instructions et sorties qui était très populaire lors de la sortie de Llama-1, car elle permettait de faire du fine-tuning d'un LLM de base pour rivaliser avec ChatGPT lui-même. Vous pouvez accéder à la version GPT-4 du jeu de données Alpaca \[ici\](https://huggingface.co/datasets/vicgalle/alpaca-gpt4.). Ci-dessous figurent quelques exemples du jeu de données : ![](https://unsloth.ai/files/60d3a95adaa157855c394401b44149e9709868a4) Vous pouvez voir qu'il y a 3 colonnes dans chaque ligne - une instruction, une entrée et une sortie. Nous combinons essentiellement chaque ligne en une grande invite comme ci-dessous. Nous l'utilisons ensuite pour ajuster finement le modèle de langage, ce qui l'a rendu très similaire à ChatGPT. Nous appelons ce processus \*\*le fine-tuning supervisé par instructions\*\*. ![](https://unsloth.ai/files/3b1f21b22edb3cea09bf8862affa370ae60e3256) \### Plusieurs colonnes pour le fine-tuning Mais un gros problème est que, pour les assistants de style ChatGPT, nous n'autorisons qu'une seule instruction / une seule invite, et non plusieurs colonnes / entrées. Par exemple, dans ChatGPT, vous pouvez voir que nous devons soumettre une seule invite, et non plusieurs invites. ![](https://unsloth.ai/files/739c0dc97ca88408acbd7b738e362bed3c63643e) Cela signifie essentiellement que nous devons « fusionner » plusieurs colonnes en une seule grande invite pour que le fine-tuning fonctionne réellement ! Par exemple, le très célèbre jeu de données Titanic comporte énormément de colonnes. Votre travail consistait à prédire si un passager avait survécu ou péri en fonction de son âge, de sa classe de passager, du prix du billet, etc. Nous ne pouvons pas simplement transmettre cela à ChatGPT ; nous devons plutôt « fusionner » ces informations en une seule grande invite. ![](https://unsloth.ai/files/645c184cd3ea4e1e395d81c03e414443f16ff8c3) Par exemple, si nous interrogeons ChatGPT avec notre unique invite « fusionnée » qui inclut toutes les informations sur ce passager, nous pouvons alors lui demander de deviner ou de prédire si le passager est mort ou a survécu. ![](https://unsloth.ai/files/d77f19d8588aaff3dbdeb5659bf2a0abb3736206) D'autres bibliothèques de fine-tuning vous obligent à préparer manuellement votre jeu de données pour le fine-tuning, en fusionnant toutes vos colonnes en une seule invite. Dans Unsloth, nous fournissons simplement la fonction appelée \`to\_sharegpt\` qui fait cela en une seule fois ! ![](https://unsloth.ai/files/6d757abba5644ffac39a50859d5a5a60b26efbe9) Maintenant, c'est un peu plus compliqué, car nous permettons beaucoup de personnalisation, mais il y a quelques points : \* Vous devez entourer toutes les colonnes d'accolades \`{}\`. Ce sont les noms de colonnes dans le vrai fichier CSV / Excel. \* Les composants textuels optionnels doivent être entourés de \`\[\[\]\]\`. Par exemple, si la colonne « input » est vide, la fonction de fusion n'affichera pas le texte et l'ignorera. Cela est utile pour les jeux de données comportant des valeurs manquantes. \* Sélectionnez la colonne de sortie ou la colonne cible / prédiction dans \`output\_column\_name\`. Pour le jeu de données Alpaca, ce sera \`output\`. Par exemple, dans le jeu de données Titanic, nous pouvons créer un grand format d'invite fusionnée comme ci-dessous, où chaque colonne / fragment de texte devient optionnel. ![](https://unsloth.ai/files/80a0095d7e1fac10111fa127c1bd2b35f8a07d0e) Par exemple, imaginez que le jeu de données ressemble à ceci avec beaucoup de données manquantes : | Embarqué | Âge | Tarif | | -------- | --- | ----- | | S | 23 | | | | 18 | 7.25 | Alors, nous ne voulons pas que le résultat soit : 1. Le passager a embarqué à S. Son âge est 23. Son tarif est \*\*VIDE\*\*. 2. Le passager a embarqué à \*\*VIDE\*\*. Son âge est 18. Son tarif est 7,25 $. Au lieu de cela, en entourant les colonnes de façon optionnelle avec \`\[\[\]\]\`, nous pouvons exclure entièrement cette information. 1. \\\[\\\[Le passager a embarqué à S.\]\] \\\[\\\[Son âge est 23.\]\] \\\[\\\[Son tarif est \*\*VIDE\*\*.\]\] 2. \\\[\\\[Le passager a embarqué à \*\*VIDE\*\*.\]\] \\\[\\\[Son âge est 18.\]\] \\\[\\\[Son tarif est 7,25 $\]\] devient : 1. Le passager a embarqué à S. Son âge est 23. 2. Son âge est 18. Son tarif est 7,25 $. ### Conversations à plusieurs tours Un problème, si vous ne l'avez pas remarqué, est que le jeu de données Alpaca est à tour unique, alors que rappelez-vous, l'utilisation de ChatGPT était interactive et vous pouviez lui parler en plusieurs tours. Par exemple, la partie gauche est ce que nous voulons, mais la partie droite, qui est le jeu de données Alpaca, ne fournit que des conversations uniques. Nous voulons que le modèle de langage ajusté apprenne d'une manière ou d'une autre à faire des conversations à plusieurs tours, tout comme ChatGPT. ![](https://unsloth.ai/files/415abaf3237603ab574a04d43eeb764b544eea5c) Nous avons donc introduit le \`conversation\_extension\` paramètre, qui sélectionne essentiellement quelques lignes aléatoires dans votre jeu de données à tour unique et les fusionne en une seule conversation ! Par exemple, si vous le réglez sur 3, nous sélectionnons aléatoirement 3 lignes et les fusionnons en 1 ! Le régler trop haut peut rendre l'entraînement plus lent, mais pourrait rendre votre chatbot et votre fine-tuning final bien meilleurs ! ![](https://unsloth.ai/files/d470e9afb208b14db6a12f51ba543cde06ac03c1) Puis définissez \`output\_column\_name\` à la colonne de prédiction / sortie. Pour le jeu de données Alpaca dataset, ce serait la colonne de sortie. Nous utilisons ensuite la \`standardize\_sharegpt\` fonction pour simplement mettre le jeu de données dans un format correct pour le fine-tuning ! Appelez toujours cette fonction ! ![](https://unsloth.ai/files/bd75ed25619b920c6932cf5f1364d50f3ef96914) \## Fine-tuning vision Le jeu de données destiné au fine-tuning d'un modèle de vision ou multimodal comprend également des entrées d'image. Par exemple, le \[Notebook vision Llama 3.2\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(11B\\)-Vision.ipynb#scrollTo=vITh0KVJ10qX) utilise un cas de radiographie pour montrer comment l'IA peut aider les professionnels de santé à analyser plus efficacement les radiographies, les scanners et les échographies. Nous utiliserons une version échantillonnée du jeu de données de radiographie ROCO. Vous pouvez accéder au jeu de données \[ici\](https://www.google.com/url?q=https%3A%2F%2Fhuggingface.co%2Fdatasets%2Funsloth%2FRadiology\_mini). Le jeu de données comprend des radiographies, des scanners et des échographies montrant des affections et des maladies médicales. Chaque image possède une légende rédigée par des experts qui la décrit. L'objectif est d'ajuster finement un VLM pour en faire un outil d'analyse utile aux professionnels de santé. Regardons le jeu de données et voyons ce que montre le premier exemple : \`\`\` Jeu de données({ features: \['image', 'image\_id', 'caption', 'cui'\], num\_rows: 1978 }) \`\`\` | Image | Légende | | --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ![](https://unsloth.ai/files/117a20db5767e07daf56898678b2129ec8c4e745) | La radiographie panoramique montre une lésion ostéolytique dans la partie postérieure droite du maxillaire avec résorption du plancher du sinus maxillaire (flèches). | Pour mettre le jeu de données en forme, toutes les tâches de fine-tuning vision doivent être formatées comme suit : \`\`\`python \[ { "role": "user", "content": \[{"type": "text", "text": instruction}, {"type": "image", "image": image} \] }, { "role": "assistant", "content": \[{"type": "text", "text": answer} \] }, \] \`\`\` Nous allons rédiger une consigne personnalisée demandant au VLM d'être un radiographe expert. Notez également qu'au lieu d'une seule instruction, vous pouvez ajouter plusieurs tours pour en faire une conversation dynamique. \`\`\`notebook-python instruction = "Vous êtes un radiographe expert. Décrivez avec précision ce que vous voyez dans cette image." def convert\_to\_conversation(sample): conversation = \[ { "role": "user", "content" : \[ {"type" : "text", "text" : instruction}, {"type" : "image", "image" : sample\["image"\]} \] }, { "role" : "assistant", "content" : \[ {"type" : "text", "text" : sample\["caption"\]} \] }, \] return { "messages" : conversation } pass \`\`\` Convertissons le jeu de données dans le « bon » format pour le fine-tuning : \`\`\`notebook-python converted\_dataset = \[convert\_to\_conversation(sample) for sample in dataset\] \`\`\` Le premier exemple est désormais structuré comme ci-dessous : \`\`\`notebook-python converted\_dataset\[0\] \`\`\` {% code overflow="wrap" %} \`\`\` {'messages': \[{'role': 'user', 'content': \[{'type': 'text', 'text': 'Vous êtes un radiographe expert. Décrivez avec précision ce que vous voyez dans cette image.'}, {'type': 'image', 'image': }\]}, {'role': 'assistant', 'content': \[{'type': 'text', 'text': 'La radiographie panoramique montre une lésion ostéolytique dans la partie postérieure droite du maxillaire avec résorption du plancher du sinus maxillaire (flèches).'}\]}\]} \`\`\` {% endcode %} Avant de faire du fine-tuning, le modèle de vision sait peut-être déjà comment analyser les images ? Vérifions si c'est le cas ! \`\`\`notebook-python FastVisionModel.for\_inference(model) # Activer pour l'inférence ! image = dataset\[0\]\["image"\] instruction = "Vous êtes un radiographe expert. Décrivez avec précision ce que vous voyez dans cette image." messages = \[ {"role": "user", "content": \[ {"type": "image"}, {"type": "text", "text": instruction} \]} \] input\_text = tokenizer.apply\_chat\_template(messages, add\_generation\_prompt = True) inputs = tokenizer( image, input\_text, add\_special\_tokens = False, return\_tensors = "pt", ).to("cuda") from transformers import TextStreamer text\_streamer = TextStreamer(tokenizer, skip\_prompt = True) \_ = model.generate(\*\*inputs, streamer = text\_streamer, max\_new\_tokens = 128, use\_cache = True, temperature = 1.5, min\_p = 0.1) \`\`\` Et le résultat : \`\`\` Cette radiographie semble être une vue panoramique de la dentition supérieure et inférieure, plus précisément un orthopantomogramme (OPG). \* La radiographie panoramique montre des structures dentaires normales. \* Il y a une zone anormale en haut à droite, représentée par une zone d'os radiotransparente, correspondant à l'antre. \*\*Observations clés\*\* \* L'os entre les dents supérieures gauche est relativement radio-opaque. \* Il y a deux grandes flèches au-dessus de l'image, suggérant la nécessité d'un examen plus approfondi de cette zone. L'une des flèches est du côté gauche, et l'autre du côté droit. Cependant, seulement \`\`\` Pour plus de détails, consultez notre section sur les jeux de données dans le \[notebook ici\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(11B\\)-Vision.ipynb#scrollTo=vITh0KVJ10qX). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/datasets-guide.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/commencer/install/windows-installation.md). # Comment fine-tuner des LLM sur Windows avec Unsloth (guide étape par étape) Vous pouvez désormais affiner des modèles directement sur votre appareil Windows local sans WSL en utilisant \[Unsloth\](https://github.com/unslothai/unsloth). Pour ce guide, il existe 3 méthodes principales que vous pouvez utiliser (\[Conda\](#method-1-windows-via-conda), \[Docker\](#method-2-docker) et \[WSL\](#method-3-wsl)).\\ Si vous avez déjà PyTorch installé sur Windows, \`pip install unsloth\` devrait fonctionner. Sinon, suivez nos guides ci-dessous : [Tutoriel Conda](https://unsloth.ai/pages/897f358231280f27c42fdd0fe983b7be0dd75dfc#method-1-windows-via-conda) [Tutoriel Docker](https://unsloth.ai/pages/897f358231280f27c42fdd0fe983b7be0dd75dfc#method-2-docker) [Tutoriel WSL](https://unsloth.ai/pages/897f358231280f27c42fdd0fe983b7be0dd75dfc#method-3-wsl) ### Unsloth Studio Nous avons lancé une nouvelle interface web appelée \[Unsloth Studio\](/docs/fr/nouveau/studio/install.md) qui fonctionne sur Windows immédiatement : \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` Utilisez la même commande pour mettre à jour. Puis, pour lancer à chaque fois : \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` Pour des instructions d'installation détaillées et les prérequis d'Unsloth Studio, \[consultez notre guide\](/docs/fr/nouveau/studio/install.md). Voici ci-dessous les instructions d’installation pour le \*\*Unsloth Core\*\*: ### Méthode n°1 - Windows via Conda : {% stepper %} {% step %} \*\*Installez Miniconda (ou Anaconda)\*\* Télécharger Anaconda \[ici\](https://www.anaconda.com/download). Nous vous suggérons d’utiliser \[Miniconda\](https://www.anaconda.com/docs/getting-started/miniconda/install#quickstart-install-instructions). Pour l’utiliser, ouvrez d’abord PowerShell - recherchez « Windows Powershell » dans le menu Démarrer : ![](https://unsloth.ai/files/90842dedf8b6468877928c76d77c08821eb58d0b) Ensuite, PowerShell s’ouvrira : ![](https://unsloth.ai/files/34d128a79a0c98eb432ec4a65dba8a2e7891b337) Ensuite, copiez-collez ce qui suit : CTRL+C, puis collez-le dans PowerShell avec CTRL+V : {% code overflow="wrap" %} \`\`\`ps Invoke-WebRequest -Uri "https://repo.anaconda.com/miniconda/Miniconda3-latest-Windows-x86\_64.exe" -OutFile ".\\miniconda.exe" Start-Process -FilePath ".\\miniconda.exe" -ArgumentList "/S" -Wait del .\\miniconda.exe \`\`\` {% endcode %} Acceptez l’avertissement et cliquez sur « Coller quand même », puis attendez. ![](https://unsloth.ai/files/325922e8d495be69f4c1eb4cb988b663640e99b2) Il télécharge l’installateur comme ci-dessous : ![](https://unsloth.ai/files/82cbafdefa1e650d620a29b9ad33053176422f91) Après l’installation, ouvrez \*\*Anaconda Powershell Prompt\*\* pour utiliser Miniconda via Démarrer -> Recherchez-le : ![](https://unsloth.ai/files/a4aaf42f51f81a1e2b71de38ad1aed5bac7553b8) Ensuite, vous verrez : ![](https://unsloth.ai/files/f52c267e3549e418f830365c9f148ca1d2905742) {% endstep %} {% step %} \*\*Créer un environnement conda\*\* \`\`\`bash conda create --name unsloth\_env python==3.12 -y conda activate unsloth\_env \`\`\` \*\*Vous verrez :\*\* ![](https://unsloth.ai/files/ae2a6f146141140ec3b4cde39388382bb5a77c68) {% endstep %} {% step %} \*\*Vérifiez \`nvidia-smi\` pour confirmer que vous avez un GPU, et recherchez la version de CUDA\*\* Après avoir tapé \`nvidia-smi\` dans PowerShell, vous devriez voir quelque chose comme ci-dessous. Si vous n’avez pas \`nvidia-smi\` ou si ce qui suit ne s’affiche pas, vous devez réinstaller \[pilotes NVIDIA\](https://www.nvidia.com/en-us/drivers/). ![](https://unsloth.ai/files/5d34640d662e5e38eb176c4569a044462a1a71d0) {% endstep %} {% step %} \*\*Installer PyTorch\*\* Lors de l’exécution de \`nvidia-smi\` vous verrez en haut à droite : « CUDA Version: 13.0 ». Installez PyTorch dans PowerShell via. Changez \`130\` pour votre version de CUDA - assurez-vous que la \[version existe\](https://pytorch.org/) et correspond à la version de votre pilote CUDA. {% code overflow="wrap" %} \`\`\`bash pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 \`\`\` {% endcode %} Vous verrez : ![](https://unsloth.ai/files/67a9df327e788af0c985dba0533ef3c055013a8e) Essayez d’exécuter ceci dans Python via \`python\` après l’installation de PyTorch : {% code overflow="wrap" %} \`\`\`python import torch print(torch.cuda.is\_available()) A = torch.ones((10, 10), device = "cuda") B = torch.ones((10, 10), device = "cuda") A @ B \`\`\` {% endcode %} Vous devriez voir une matrice de 10. Vérifiez aussi que la première valeur est True. ![](https://unsloth.ai/files/20f92487f8005b9f55dd113008f28208c96406c6) {% endstep %} {% step %} \*\*Installez Unsloth (uniquement si PyTorch fonctionne !)\*\* {% hint style="danger" %} \*\*Confirmez que PyTorch fonctionne correctement et s’exécute - sinon PyTorch est cassé et cela signifie malheureusement que votre machine Windows peut avoir besoin d’une réinstallation des pilotes CUDA.\*\* {% endhint %} Dans PowerShell (après être sorti de Python via \`exit()\` , faites cela et attendez : \`\`\`bash pip install unsloth \`\`\` {% endstep %} {% step %} \*\*Vérifiez qu’Unsloth fonctionne\*\* Utilisez maintenant n’importe quel script dans \[Notebooks Unsloth\](/docs/fr/commencer/unsloth-notebooks.md) (enregistrez-le dans un fichier .py), ou utilisez le script de base ci-dessous : {% code expandable="true" %} \`\`\`python from unsloth import FastLanguageModel, FastModel import torch from trl import SFTTrainer, SFTConfig from datasets import load\_dataset max\_seq\_length = 512 url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified\_chip2.jsonl" dataset = load\_dataset("json", data\_files = {"train" : url}, split = "train") model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/gemma-3-270m-it", max\_seq\_length = max\_seq\_length, # Choisissez-en n’importe quelle pour un contexte long ! load\_in\_4bit = True, # Quantification en 4 bits. False = LoRA en 16 bits. load\_in\_8bit = False, # Quantification en 8 bits load\_in\_16bit = False, # LoRA en 16 bits full\_finetuning = False, # À utiliser pour un fine-tuning complet. trust\_remote\_code = False, # Activez pour prendre en charge les nouveaux modèles # token = "hf\_...", # en utilisez un si vous utilisez des modèles protégés ) # Applique le patch au modèle et ajoute des poids LoRA rapides model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\], lora\_alpha = 16, lora\_dropout = 0, # Prend en charge n’importe quelle valeur, mais = 0 est optimisé bias = "none", # Prend en charge n’importe quelle valeur, mais = "none" est optimisé # \[NOUVEAU\] "unsloth" utilise 30 % de VRAM en moins, permet des tailles de lot 2x plus grandes ! use\_gradient\_checkpointing = "unsloth", # True ou "unsloth" pour un contexte très long random\_state = 3407, max\_seq\_length = max\_seq\_length, use\_rslora = False, # Nous prenons en charge LoRA avec stabilisation du rang loftq\_config = None, # Et LoftQ ) trainer = SFTTrainer( model = model, train\_dataset = dataset, tokenizer = tokenizer, args = SFTConfig( max\_seq\_length = max\_seq\_length, per\_device\_train\_batch\_size = 2, gradient\_accumulation\_steps = 4, warmup\_steps = 10, max\_steps = 60, logging\_steps = 1, output\_dir = "outputs", optim = "adamw\_8bit", seed = 3407, dataset\_num\_proc = 1, ), ) trainer.train() \`\`\` {% endcode %} Vous devriez voir : \`\`\`bash 🦥 Unsloth : va patcher votre ordinateur pour activer un fine-tuning gratuit 2x plus rapide. 🦥 Unsloth Zoo va maintenant tout patcher pour rendre l’entraînement plus rapide ! ==((====))== Unsloth 2026.1.4 : patching rapide de Gemma3. Transformers : 4.57.6. \\\\ /| NVIDIA GeForce RTX 3060. Nombre de GPU = 1. Mémoire max : 12.0 Go. Plateforme : Windows. O^O/ \\\_/ \\ Torch : 2.10.0+cu130. CUDA : 8.6. Kit d’outils CUDA : 13.0. Triton : 3.6.0 \\ / Bfloat16 = TRUE. FA \[Xformers = 0.0.34. FA2 = False\] "-\_\_\_\_-" Licence gratuite : http://github.com/unslothai/unsloth Unsloth : le téléchargement rapide est activé - ignorez les barres de téléchargement rouges ! Unsloth : Gemma3 ne prend pas en charge SDPA - bascule vers fast eager. Unsloth : mise en place de \`model.base\_model.model.model\` pour exiger des gradients Unsloth : Tokenizing \["text"\] (num\_proc=1): 0%| | 0/210289 \[00:00![](https://unsloth.ai/files/854cd18a3e42dc573da86030f2e5cebb2567ab8b)\ \ {% endstep %} {% endstepper %} ### Méthode n°2 - Docker : Docker est peut-être le moyen le plus simple pour les utilisateurs de Windows de commencer avec Unsloth, car aucune configuration ni aucun problème de dépendance n’est nécessaire. \[\*\*\`unsloth/unsloth\`\*\*\](https://hub.docker.com/r/unsloth/unsloth) est la seule image Docker d’Unsloth. Pour \[Blackwell\](/docs/fr/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md) et les GPU de série 50, utilisez cette même image - aucune image séparée n’est nécessaire. Pour les instructions d’installation, veuillez suivre notre \[guide Docker\](/docs/fr/blog/how-to-fine-tune-llms-with-unsloth-and-docker.md), sinon voici un guide de démarrage rapide : {% stepper %} {% step %} \*\*Installez Docker et NVIDIA Container Toolkit.\*\* Installez Docker via \[Linux\](https://docs.docker.com/engine/install/) ou \[Desktop\](https://docs.docker.com/desktop/) (autre). Puis installez \[NVIDIA Container Toolkit\](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#installation): \`\`\`bash export NVIDIA\_CONTAINER\_TOOLKIT\_VERSION=1.17.8-1 sudo apt-get update && sudo apt-get install -y \\\\ nvidia-container-toolkit=${NVIDIA\_CONTAINER\_TOOLKIT\_VERSION} \\\\ nvidia-container-toolkit-base=${NVIDIA\_CONTAINER\_TOOLKIT\_VERSION} \\\\ libnvidia-container-tools=${NVIDIA\_CONTAINER\_TOOLKIT\_VERSION} \\\\ libnvidia-container1=${NVIDIA\_CONTAINER\_TOOLKIT\_VERSION} \`\`\` {% endstep %} {% step %} \*\*Exécutez le conteneur.\*\* \[\*\*\`unsloth/unsloth\`\*\*\](https://hub.docker.com/r/unsloth/unsloth) est la seule image Docker d’Unsloth. \`\`\`bash docker run -d -e JUPYTER\_PASSWORD="mypassword" \\ -p 8888:8888 -p 2222:22 \\\\ -v $(pwd)/work:/workspace/work \\ --gpus all \\ unsloth/unsloth \`\`\` {% endstep %} {% step %} \*\*Accédez à Jupyter Lab\*\* Allez dans \[http://localhost:8888\](http://localhost:8888/) et ouvrez Unsloth. Accédez aux \`notebooks-unsloth\` onglets pour voir les notebooks Unsloth. {% endstep %} {% step %} \*\*Commencez à entraîner avec Unsloth\*\* Si vous débutez, suivez notre guide \[Guide de fine-tuning\](/docs/fr/commencer/fine-tuning-llms-guide.md), \[Guide RL\](/docs/fr/commencer/reinforcement-learning-rl-guide.md) ou enregistrez/copiez simplement l’un de nos \[notebooks\](/docs/fr/commencer/unsloth-notebooks.md). {% endstep %} {% step %} \*\*Problèmes Docker - GPU non détecté ?\*\* Essayez de faire du WSL via \[#method-2-wsl\](#method-2-wsl "mention") {% endstep %} {% endstepper %} ### Méthode n°3 - WSL : {% stepper %} {% step %} \*\*Installez WSL\*\* Ouvrez l’Invite de commandes, le Terminal, et installez Ubuntu. Définissez le mot de passe si demandé. \`\`\`bash wsl.exe --install Ubuntu-24.04 wsl.exe -d Ubuntu-24.04 \`\`\` {% endstep %} {% step %} \*\*Si vous n’avez PAS fait (1), donc vous avez déjà installé WSL\*\*\*\*, entrez dans WSL en tapant \`wsl\` et ENTRÉE dans l’invite de commandes\*\* \`\`\`bash wsl \`\`\` {% endstep %} {% step %} \*\*Installez Python\*\* {% code overflow="wrap" %} \`\`\`bash sudo apt update sudo apt install python3 python3-full python3-pip python3-venv -y \`\`\` {% endcode %} {% endstep %} {% step %} \*\*Installer PyTorch\*\* {% code overflow="wrap" %} \`\`\`bash pip install torch torchvision --force-reinstall --index-url https://download.pytorch.org/whl/cu130 \`\`\` {% endcode %} Si vous rencontrez des problèmes d’autorisation, utilisez \`–break-system-packages\` afin que \`pip install torch torchvision --force-reinstall --index-url https://download.pytorch.org/whl/cu130 –break-system-packages\` {% endstep %} {% step %} \*\*Installez Unsloth et Jupyter Notebook\*\* {% code overflow="wrap" %} \`\`\`bash pip install unsloth jupyter \`\`\` {% endcode %} Si vous rencontrez des problèmes d’autorisation, utilisez \`–-break-system-packages\` afin que \`pip install unsloth jupyter –-break-system-packages\` {% endstep %} {% step %} \*\*Lancez Unsloth via Jupyter Notebook\*\* {% code overflow="wrap" %} \`\`\`bash jupyter notebook \`\`\` {% endcode %} Puis ouvrez nos notebooks dans \[Notebooks Unsloth\](/docs/fr/commencer/unsloth-notebooks.md)et chargez-les ! Vous pouvez aussi aller dans les notebooks Colab et télécharger > télécharger .ipynb puis les charger. !\[\](/files/a724ef997c60e3aee78efcd8d9e3ac4458412bba) {% endstep %} {% endstepper %} {% hint style="warning" %} Si vous utilisez GRPO ou prévoyez d’utiliser vLLM, à l’heure actuelle vLLM ne prend pas en charge Windows directement, mais uniquement via WSL ou Linux. {% endhint %} ### \*\*Dépannage /\*\* Avancé Pour \*\*instructions d’installation avancées\*\* ou si vous voyez des erreurs étranges pendant les installations : 1. Installez \`torch\` et \`triton\`. Allez sur pour l’installer. Par exemple \`pip install torch torchvision torchaudio triton\` 2. Vérifiez si CUDA est correctement installé. Essayez \`nvcc\`. Si cela échoue, vous devez installer \`cudatoolkit\` ou les pilotes CUDA. 3. Si vous utilisez un GPU Intel, vous devrez suivre notre \[guide Windows Intel\](/docs/fr/commencer/install/intel.md#windows-only-runtime-configurations) 4. Installez \`xformers\` manuellement. Vous pouvez essayer d’installer \`vllm\` et voir si \`vllm\` réussit. Vérifiez si \`xformers\` a réussi avec \`python -m xformers.info\` Allez sur . Une autre option est d’installer \`flash-attn\` pour les GPU Ampere. 5. Vérifiez bien que vos versions de Python, CUDA, CUDNN, \`torch\`, \`triton\`et \`xformers\` sont compatibles entre elles. Le \[tableau de compatibilité PyTorch\](https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix) peut être utile. 6. Enfin, installez \`bitsandbytes\` et vérifiez-le avec \`python -m bitsandbytes\` 7. Si Unsloth ne détecte pas ou n’utilise pas votre GPU et si vous utilisez notre conteneur Docker sur Windows, votre version du kit d’outils CUDA \`nvcc --version\` doit correspondre à la version de CUDA affichée par nvidia-smi sur le GPU hôte. La prise en charge des conteneurs Docker sur Windows n’est pas automatique. \[Vous devez suivre le guide de Docker\](https://docs.docker.com/desktop/features/gpu/). ### Désinstallation La manière recommandée de supprimer complètement Unsloth Studio est d'utiliser le script de désinstallation correspondant à votre OS. Il arrête tous les serveurs en cours d'exécution, supprime l'application, la commande CLI, les données du lanceur, les raccourcis et les entrées spécifiques à la plateforme (macOS \`.app\` bundle + Launch Services ; menu Démarrer Windows + registre + PATH) : \`\`\`ps1 irm https://raw.githubusercontent.com/unslothai/unsloth/main/scripts/uninstall.ps1 | iex \`\`\` #### Désinstallation manuelle Si vous préférez supprimer uniquement certaines parties : \*\*1. Supprimer uniquement l'application\*\* (conserve l'historique, les discussions, les points de contrôle et les exportations intacts) : \* \`Remove-Item -Recurse -Force "$HOME\\.unsloth\\studio\\unsloth\_studio"\` \*\*2. Supprimer complètement Unsloth\*\* (conserve intacts les autres outils Unsloth) : \* \`Remove-Item -Recurse -Force "$HOME\\.unsloth\\studio"\` \*\*3. Supprimer tout ce qui est lié à Unsloth :\*\* \* \`Remove-Item -Recurse -Force "$HOME\\.unsloth"\` {% hint style="warning" %} Remarque : l'étape 3 supprime tout l'historique, les discussions, les points de contrôle du modèle et les exportations. Cela ne peut pas être annulé. {% endhint %} \*\*4. Supprimer les raccourcis et les liens symboliques :\*\* \`\`\`shellscript Remove-Item -Force "$HOME\\Desktop\\Unsloth Studio.lnk" Remove-Item -Force "$env:APPDATA\\Microsoft\\Windows\\Start Menu\\Programs\\Unsloth Studio.lnk" \`\`\` \*\*5. Supprimer la commande CLI :\*\* \* \*\*Windows (PowerShell) :\*\* Le programme d'installation a ajouté le \`Scripts\` répertoire au PATH de votre utilisateur. Pour le supprimer, ouvrez Paramètres → Système → À propos → Paramètres système avancés → Variables d'environnement, puis recherchez \`Path\` sous Variables utilisateur, et supprimez l'entrée pointant vers \`.unsloth\\studio\\...\\Scripts\`. {% hint style="info" %} Remarque : les étapes 1 à 5 ne touchent pas aux fichiers de modèle HF que vous avez téléchargés. Voir la section Suppression des fichiers de modèle HF mis en cache ci-dessous si vous souhaitez récupérer cet espace. {% endhint %} ### \*\*Suppression des fichiers de modèle HF mis en cache\*\* Vous pouvez supprimer d'anciens fichiers de modèle soit depuis l'icône de corbeille dans la recherche de modèles, soit en supprimant le dossier de modèle mis en cache correspondant dans le répertoire de cache Hugging Face par défaut. Par défaut, Hugging Face utilise \`~/.cache/huggingface/hub/\` sur macOS/Linux/WSL et \`C:\\Users\\\\.cache\\huggingface\\hub\\\` sur Windows. \* \*\*Windows :\*\* \`%USERPROFILE%\\.cache\\huggingface\\hub\\\` Si \`HF\_HUB\_CACHE\` ou \`HF\_HOME\` est défini, utilisez cet emplacement à la place. Sur Linux et WSL, \`XDG\_CACHE\_HOME\` peut également modifier la racine de cache par défaut. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/commencer/install/windows-installation.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/lm-studio.md). # Deploying models to LM Studio You can run and deploy your fine-tuned LLM directly in LM Studio. \[LM Studio\](https://lmstudio.ai/) enables easy running and deployment of \*\*GGUF\*\* models (llama.cpp format). You can use our \[LM Studio notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/FunctionGemma\_\\(270M\\)-LMStudio.ipynb) or follow the instructions below: 1. \*\*Export your Unsloth fine-tuned model to \`.gguf\`\*\* 2. \*\*Import / download the GGUF into LM Studio\*\* 3. \*\*Load it in Chat\*\* (or run it behind an OpenAI-compatible local API) ![](https://unsloth.ai/files/Wbf0LEC8HXMpqIPKFZYa) Before fine-tuning in LM Studio ![](https://unsloth.ai/files/SZ2ctpKOXmOn0ahCrPRM) After fine-tuning in LM Studio \### 1) Export to GGUF (from Unsloth) If you already exported a \`.gguf\`, skip to \*\*Importing into LM Studio\*\*. \`\`\`python # Save locally (creates GGUF artifacts in the folder) model.save\_pretrained\_gguf("my\_model\_gguf", tokenizer, quantization\_method = "q4\_k\_m") # model.save\_pretrained\_gguf("my\_model\_gguf", tokenizer, quantization\_method = "q8\_0") # model.save\_pretrained\_gguf("my\_model\_gguf", tokenizer, quantization\_method = "f16") # Or push GGUF to the Hugging Face Hub model.push\_to\_hub\_gguf("hf\_username/my\_model\_gguf", tokenizer, quantization\_method = "q4\_k\_m") \`\`\` {% hint style="info" %} \`q4\_k\_m\` is usually the default for local runs. \`q8\_0\` is the optimum for near full precision quality. \`f16\` is largest / slowest, but original unquantized precision. {% endhint %} ### 2) Import the GGUF into LM Studio {% tabs %} {% tab title="CLI Import (lms import)" %} LM Studio provides a CLI called \`lms\` that can import a local \`.gguf\` into LM Studio’s models folder. \*\*Import a GGUF file:\*\* \`\`\`bash lms import /path/to/model.gguf \`\`\` \*\*Keep the original file (copy instead of move):\*\* \`\`\`bash lms import /path/to/model.gguf --copy \`\`\` **Click for more customizable private settings** \*\*Keep the model where it is (symlink):\*\* This is helpful for large models stored on a dedicated drive. \`\`\`bash lms import /path/to/model.gguf --symbolic-link \`\`\` \*\*Skip prompts and choose the target namespace yourself:\*\* \`\`\`bash lms import /path/to/model.gguf --user-repo my-user/my-finetuned-models \`\`\` \*\*Dry-run (shows what will happen):\*\* \`\`\`bash lms import /path/to/model.gguf --dry-run \`\`\` After importing, the model should appear in LM Studio under \*\*My Models\*\*. ![](https://unsloth.ai/files/dY4UqwcuLsYbVQAw5ebU) {% endtab %} {% tab title="From Hugging Face" %} If you pushed your GGUF repo to Hugging Face, you can download it directly from within LM Studio. \*\*Option A: Use LM Studio’s in-app downloader\*\* 1. Open LM Studio 2. Go to the \*\*Discover\*\* tab 3. Search for \`hf\_username/repo\_name\` (or paste the Hugging Face URL) 4. Download the quant you want (e.g. \`Q4\_K\_M\`) \*\*Option B: Use the CLI downloader\*\* \`\`\`bash # Download from HF by repo name lms get hf\_username/my\_model\_gguf # Pick a quantization with @ lms get hf\_username/my\_model\_gguf@Q4\_K\_M \`\`\` {% endtab %} {% tab title="Manual Import (folder structure)" %} If you don’t want to use the CLI, you can place the \`.gguf\` file into LM Studio’s expected model directory structure. LM Studio expects models to look like this: \`\`\` ~/.lmstudio/models/ └── publisher/ └── model/ └── model-file.gguf \`\`\` Example: \`\`\` ~/.lmstudio/models/ └── my-name/ └── my-finetune/ └── my-finetune-Q4\_K\_M.gguf \`\`\` Then open LM Studio and check \*\*My Models\*\*. \*\*Tip:\*\* You can manage / verify your models directory from the \*\*My Models\*\* tab in LM Studio. {% endtab %} {% endtabs %} ### 3) Load and chat in LM Studio 1. Open LM Studio → \*\*Chat\*\* 2. Open the \*\*model loader\*\* 3. Select your imported model 4. (Optional) adjust load settings (GPU offload, context length, etc.) 5. Chat normally in the UI ### 4) Serve your fine-tuned model as a local API (OpenAI-compatible) LM Studio can serve your loaded model behind an OpenAI-compatible API (handy for apps like Open WebUI, custom agents, scripts, etc.). {% tabs %} {% tab title="GUI (Developer tab)" %} 1. Load your model in LM Studio 2. Go to the \*\*Developer\*\* tab 3. Start the local server 4. Use the shown base URL (default is typically \`http://localhost:1234/v1\`) {% endtab %} {% tab title="CLI (lms load + lms server start)" %} #### 1) List available models \`\`\`bash lms ls \`\`\` #### 2) Load your model (optional flags) \`\`\`bash lms load \--gpu=auto --context-length=8192 \`\`\` Notes: \* \`--gpu=1.0\` means “try to offload 100% to GPU” \* You can set a stable identifier: \`\`\`bash lms load \--identifier="my-finetuned-model" \`\`\` #### 3) Start the server \`\`\`bash lms server start --port 1234 \`\`\` {% endtab %} {% endtabs %} \*\*Quick test: list models\*\* \`\`\`bash curl http://localhost:1234/v1/models \`\`\` \*\*Python example (OpenAI SDK):\*\* {% code expandable="true" %} \`\`\`python from openai import OpenAI client = OpenAI( base\_url="http://localhost:1234/v1", api\_key="lm-studio", # LM Studio may not require a real key; this is a common placeholder ) resp = client.chat.completions.create( model="model-identifier-from-lm-studio", messages=\[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Hello! What did I fine-tune you to do?"}, \], temperature=0.7, # adjust temperature according to your model needs ) print(resp.choices\[0\].message.content) \`\`\` {% endcode %} \*\*cURL example (chat completions):\*\* {% code expandable="true" %} \`\`\`bash curl http://localhost:1234/v1/chat/completions \\ -H "Content-Type: application/json" \\ -d '{ "model": "model-identifier-from-lm-studio", "messages": \[ {"role": "user", "content": "Say this is a test!"} \], "temperature": 0.7 # adjust temperature according to your model needs }' \`\`\` {% endcode %} {% hint style="info" %} \*\*Debugging tip:\*\* If you’re troubleshooting formatting/templates, you can inspect the \*raw\* prompt LM Studio sends to the model by running: \`lms log stream\` {% endhint %} ### Troubleshooting #### \*\*Model runs in Unsloth, but LM Studio output is gibberish / repeats\*\* This is almost always a \*\*prompt template / chat template mismatch\*\*. LM Studio will \*\*auto-detect\*\* the prompt template from the GGUF metadata when possible, but custom or incorrectly-tagged models may need a manual override. \*\*Fix:\*\* 1. Go to \*\*My Models\*\* → click the gear ⚙️ next to your model 2. Find \*\*Prompt Template\*\* and set it to match the template you trained with 3. Alternatively, in the Chat sidebar: enable the \*\*Prompt Template\*\* box (you can force it to always show) #### LM Studio doesn’t show my model in “My Models” \* Prefer \`lms import /path/to/model.gguf\` \* Or confirm the file is in the correct folder structure: \`~/.lmstudio/models/publisher/model/model-file.gguf\` #### OOM / slow performance \* Use a smaller quant (ex: \`Q4\_K\_M\`) \* Reduce context length \* Adjust GPU offload (LM Studio “Per-model defaults” / load settings) \*\*\* ### More resources \* \[LM Studio + Unsloth blog post\](https://lmstudio.ai/blog/functiongemma-unsloth) (FunctionGemma walkthrough): \* LM Studio \[Import Models docs\](https://lmstudio.ai/docs/app/advanced/import-model) \* LM Studio \[Prompt Template docs\](https://lmstudio.ai/docs/app/advanced/prompt-template) \* LM Studio \[OpenAI-compatible API docs\](https://lmstudio.ai/docs/developer/openai-compat) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/lm-studio.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/nouveau/studio/data-recipe.md). # Recettes de données Unsloth Data Recipes d’Unsloth Studio vous permet de téléverser des documents comme des PDF ou des fichiers CSV et de les transformer en jeux de données utilisables / synthétiques. Créez et modifiez des jeux de données visuellement via un flux de travail en graphe de nœuds. Ce guide vous initiera aux bases avant que vous ne plongiez dans Unsloth Data Recipes. ![](https://unsloth.ai/files/cdbdb9ee45f65c5b9d2621ee85677f55f5ff4b08) \### Comment fonctionne Data Recipes Data Recipes suit le même cheminement de base. Vous ouvrez la page des recettes, créez ou choisissez une recette, construisez le flux de travail dans l’éditeur, le validez, lancez un aperçu, puis exécutez le jeu de données complet une fois que le résultat semble correct. Ajoutez des données de départ et des blocs de génération, validez le flux de travail, prévisualisez un échantillon de sortie, puis lancez une génération complète du jeu de données. Unsloth Data Recipes est propulsé par \*\*NVIDIA Nemo\*\* \[\*\*Data Designer\*\*\](https://github.com/NVIDIA-NeMo/DataDesigner). ![](https://unsloth.ai/files/180bd7cd3cbd0cc95559d28cb19d6385f8c94f50) Exemple de génération de jeu de données et de fine-tuning d’un modèle À première vue, un flux de travail habituel devrait ressembler à ceci : 1. Ouvrez la page des recettes. 2. Créez une nouvelle recette ou ouvrez-en une existante. 3. Ajoutez des blocs pour définir le flux de travail de votre jeu de données. 4. Cliquez sur \*\*Validez\*\* pour détecter rapidement les problèmes de configuration. 5. Lancez un aperçu pour examiner rapidement des lignes d’exemple. 6. Lancez une génération complète du jeu de données lorsque la recette est prête. 7. Suivez la progression et la sortie en direct dans le graphe ou dans \*\*Exécutions\*\* vue pour plus de détails. 8. Sélectionnez le jeu de données résultant dans \*\*Unsloth\*\* et affinez un modèle. ### Commencer La page des recettes est le point d’entrée principal. Les recettes sont stockées localement dans le navigateur, ce qui vous permet de retrouver plus tard votre travail enregistré. À partir de là, vous pouvez créer une recette vierge ou ouvrir une recette d’apprentissage guidée. {% hint style="info" %} Les recettes peuvent être exportées et importées, il est donc facile de partager des flux de travail avec d’autres utilisateurs d’Unsloth :tada:. Si vous essayez de créer un schéma de jeu de données spécifique, demandez sur le Discord d’Unsloth. Quelqu’un a peut-être déjà une recette à partager. {% endhint %} ![](https://unsloth.ai/files/7c4db81edc56c672f9d4642a0a9dd85b863389de) Page d’accueil des recettes Si vous êtes nouveau dans le concept des flux de travail, les recettes d’apprentissage sont le moyen le plus rapide de voir comment les données de départ, les prompts, les expressions et les validateurs s’assemblent dans un exemple fonctionnel. Si vous connaissez déjà la forme du jeu de données que vous souhaitez, partir de zéro est généralement plus rapide. #### Choisissez un point de départ | Si vous souhaitez : | Commencez avec : | | | --- | --- | --- | | **Construire rapidement un flux de travail personnalisé** | **Commencer à vide** | | | **Apprendre le produit à partir d’un exemple** | **Commencer à partir d’une recette d’apprentissage** | | | **Poursuivre le travail précédent** | **Ouvrir une recette enregistrée** | | \### Ce que vous construisez dans l’éditeur L’éditeur est l’endroit où la recette prend forme. Vous ajoutez des blocs depuis la palette de blocs, les configurez dans des boîtes de dialogue, les connectez sur le canevas, puis validez ou exécutez le flux de travail. ![](https://unsloth.ai/files/aee39d566df13a802d5344c4030bb11523d69459) Exemple de construction d’un flux de travail de description de produit {% columns %} {% column %} L’éditeur comporte quelques parties essentielles : \* L’en-tête de la recette, où vous renommez la recette et basculez entre \*\*Éditeur\*\* et \*\*Exécutions\*\* \* Le canevas, où le graphe de la recette est affiché \* La palette de blocs, où vous ajoutez de nouveaux blocs \* Les boîtes de dialogue de configuration, où vous définissez les prompts, les références, les alias de modèle, les validateurs et les paramètres de seed. \* Les \*\*Exécutez\*\* et \*\*Validez\*\* contrôles flottants \* il faut ajouter davantage ici {% endcolumn %} {% column %} Les blocs les plus courants dans reciper sont : \* \*\*Seed\*\* pour les données d’entrée provenant de Hugging Face, de fichiers structurés locaux (ou de documents non structurés qui sont découpés en lignes). \* \*\*LLM + Modèles\*\* pour les fournisseurs, les configurations de modèle, les blocs de génération LLM et les profils d’outils partagés. \* \*\*Expression\*\* pour les transformations basées sur Jinja2 qui ne nécessitent pas d’appel LLM. \* \*\*Validateurs\*\* pour filtrer le code généré incorrect grâce à des linters intégrés pour Python, SQL et JavaScript/TypeScript. \* \*\*Échantillonneurs\*\* pour les colonnes déterministes telles que les catégories et sous-catégories. {% endcolumn %} {% endcolumns %} ### Comment fonctionnent les références La plupart des blocs qui produisent des données (avec quelques exceptions) deviennent une référence pour les blocs suivants. C’est l’une des idées principales derrière Data Recipes. Vous créez une valeur une seule fois, puis vous la réutilisez dans les prompts, les expressions, les sorties structurées et les étapes de validation. {% hint style="info" %} Les expressions Jinja vous aident à travailler avec des valeurs qui existent déjà dans la recette. Vous pouvez référencer des champs imbriqués comme \`{{customer.first\_name}}\` , associer des valeurs comme \`{{customer.first\_name}} {{customer.last\_name}}\` et ajouter une logique conditionnelle avec des motifs tels que \`{% if condition %}...{% endif %}\` {% endhint %} ![](https://unsloth.ai/files/7b5cb51697e84cfd290b2ca4b58eacfc825a81f1) Exemple de références affichées dans l’éditeur Par exemple : \* Un bloc de catégorie nommé \`domain\` peut être référencé comme \`{{ domain }}\` \* une colonne de seed peut être utilisée directement dans un prompt LLM, les colonnes de vos données de départ (par ex. colonnes d’un jeu de données HF, csv) \* une sortie LLM structurée peut exposer des champs pour des prompts ultérieurs \* un bloc d’expression peut combiner des valeurs antérieures sans autre appel au modèle ### Que se passe-t-il ensuite ? Les exécutions d’aperçu servent à itérer rapidement. Elles renvoient des lignes d’exemple et une analyse dans l’éditeur afin que vous puissiez inspecter les données générées avant de lancer une exécution complète. Les exécutions complètes créent un artefact de jeu de données local persistant. Cette sortie apparaît ensuite dans le sélecteur de jeux de données local d’Unsloth, où vous pouvez la réexaminer et l’utiliser pour le fine-tuning. Vous pouvez éventuellement publier votre jeu de données sur votre dépôt Hugging Face. ### Blocs de construction principaux {% columns %} {% column %} ![](https://unsloth.ai/files/7b0e60f03b4463cbdc4c54a2de37bea4d84a4fad) Blocs de construction principaux {% endcolumn %} {% column %} ![](https://unsloth.ai/files/bbdd0c8ab7539ce56a2e1023f8f55ec9277d7f92) Blocs de modèle et LLM {% endcolumn %} {% endcolumns %} #### La configuration du modèle est divisée en deux couches utilisables : \* \*\*Fournisseur de modèle\*\* définit le point de terminaison et l’authentification \* \*\*Configuration du modèle\*\* définit le nom du modèle et les paramètres d’inférence Cette configuration fonctionne avec des fournisseurs hébergés, des points de terminaison auto-hébergés, \`vLLM\` , \`llama.cpp\` , ou toute API compatible OpenAI que vous exécutez en dehors d’Unsloth. {% hint style="info" %} Les recettes ne sont pas limitées à un seul modèle. Vous pouvez ajouter plusieurs \*\*Fournisseurs de modèles\*\* et \*\*Configuration du modèle\*\* blocs, puis utiliser différents modèles pour différentes étapes, par exemple un pour le code et un autre pour les tâches de texte générales. {% endhint %} Après la configuration du modèle, vous pouvez utiliser quatre types de blocs LLM : | Bloc | Sortie | Idéal pour | | ------------- | ---------------- | -------------------------------------------------------------------------- | | LLM Texte | Texte libre | Instructions, explications, conversations et descriptions | | LLM Structuré | JSON | Sortie nécessitant des champs fixes et une structure prévisible | | LLM Code | Code | Python, SQL, Typescript et autres tâches de génération de code | | LLM Juge | Évaluation notée | Notation des sorties avec un ou plusieurs scores définis par l’utilisateur | #### Profils d’outils {% columns %} {% column %} Les blocs de profil d’outil définissent un accès partagé aux outils basé sur MCP pour un ou plusieurs blocs LLM. Utilisez-les lorsqu’une étape de génération a besoin d’outils, comme pour consulter la documentation du code via \`Context7\`. L’image à gauche montre Context7 MCP ajouté et configuré dans la boîte de dialogue du bloc Profil d’outil : {% endcolumn %} {% column %} ![](https://unsloth.ai/files/52328c78b475776be005b9268ef294cd47195ac3) {% endcolumn %} {% endcolumns %} #### Validateurs {% columns %} {% column %} Le bloc Validor cible principalement le bloc de code LLM en exécutant le code généré à travers un linter et une validation de syntaxe ; cela vous aide à garder les lignes de code mauvaises ou invalides hors du jeu de données final en les filtrant. Les options intégrées couvrent la validation de Python, SQL et JavaScript/TypeScript. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/7630a18317ebab81f339a3fa709cacf3599b0ee8) {% endcolumn %} {% endcolumns %} ### Valider, prévisualiser et exécuter Une fois que le flux de travail de la recette est en place, l’étape suivante est l’exécution. Le schéma recommandé est : validez d’abord, prévisualisez pour un retour rapide et inspectez les données générées dans la vue des exécutions, puis exécutez le jeu de données complet lorsque vous estimez que la sortie correspond à votre plan. Utilisez les contrôles d’exécution dans l’ordre suivant : {% stepper %} {% step %} #### Validez Cliquez sur \*\*Validez\*\* pour détecter les problèmes de configuration. {% endstep %} {% step %} #### Aperçu Lancez un aperçu pour examiner les lignes d’exemple et l’analyse {% endstep %} {% step %} #### Affiner Affinez les prompts, les références, les paramètres de seed ou les validateurs. Itérez jusqu’à ce que les données générées vous satisfassent {% endstep %} {% step %} #### Lancer la génération complète du jeu de données {% endstep %} {% endstepper %} ![](https://unsloth.ai/files/cfbc9edadecc652b83350ba5e89150e56d8ffc28) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/nouveau/studio/data-recipe.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/vllm.md). # Connect vLLM to Unsloth for Local Chat Inference Learn how to connect \*\*vLLM to\*\* \[\*\*Unsloth\*\*\](https://github.com/unslothai/unsloth) using vLLM’s \*\*OpenAI-compatible API\*\* so you can serve models and chat with them locally inside a open-source UI chat interface. This guide walks through installing vLLM, launching a local vLLM server, configuring the API base URL, loading available model IDs, and selecting your hosted vLLM model. By the end, your vLLM-served models will appear alongside local models, giving you a fast and flexible way to run external LLM inference from a UI chat interface. ### Setup {% stepper %} {% step %} #### Install vLLM Install vLLM first so you can run the \`vllm serve\` command. Follow the official \[vLLM install guide\](https://docs.vllm.ai/en/stable/getting\_started/installation/) for your platform and hardware. After installing, check that vLLM works in your terminal: \`vllm --help\` {% endstep %} {% step %} #### Choose a model vLLM serves models from Hugging Face. For example, start a vLLM server with an Unsloth model: \`\`\`bash vllm serve unsloth/gemma-4-26B-A4B-it \\ --dtype auto \`\`\` This exposes an API endpoint at: \`http://localhost:8000/v1\` To require an API key, add: \`\`\`bash --api-key token-abc123 \`\`\` {% endstep %} {% step %} #### Connect vLLM to Unsloth Open \*\*Settings → Connections\*\*, then click \*\*Add Connection\*\*. Select \*\*vLLM\*\*, then enter your server details. ![](https://unsloth.ai/files/BdMZR4w9uCCFMuVkKflR) Enter your vLLM server details: \* \*\*API key:\*\* leave empty unless you started vLLM with --api-key \* \*\*Base URL:\*\* for example, \* \*\*Reasoning model:\*\* enable this if the served model supports thinking \* \*\*Model IDs:\*\* click \*\*Load Models\*\*, or enter custom IDs manually After you click \*\*Add Connection\*\*, the models you enabled will appear under \*\*Connection\*\* in the model dropdown. {% endstep %} {% step %} #### Ready to Chat After saving the connection, your vLLM model will appear under \*\*Connected\*\* in the model dropdown. Select it to start chatting through your vLLM server. ![](https://unsloth.ai/files/ocXVg6g06AAJQrmgyze8) {% hint style="info" %} If your vLLM server is slow to respond (especially during model loading), you can adjust the timeout: \`\`\`bash AIOHTTP\_CLIENT\_TIMEOUT\_MODEL\_LIST=30 \`\`\` {% endhint %} {% endstep %} {% endstepper %} ### Common vLLM arguments The example above uses the core serving settings. You can add more vllm serve arguments depending on your model and hardware. Common options include: \`\`\`bash vllm serve unsloth/gemma-4-26B-A4B-it \\ --dtype auto \\ --host 0.0.0.0 \\ --port 8000 \\ --api-key token-abc123 \\ --max-model-len 8192 \\ --gpu-memory-utilization 0.9 \`\`\` For the full list of vLLM server arguments, see the official vLLM \[OpenAI-compatible server\](https://docs.vllm.ai/en/stable/serving/openai\_compatible\_server/) docs. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/vllm.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/openclaw.md). # How to Run Local AI Models with OpenClaw This guide will show how to use open LLMs locally with \*\*OpenClaw by connecting it to Unsloth\*\*. OpenClaw is an \*\*open-source AI agent\*\* interface that connects to a model to run tasks across your project. OpenClaw is able to works with any local model by connecting through \*\*Unsloth’s OpenAI-compatible API\*\*: including DeepSeek, Qwen, Gemma, and more. OpenClaw acts as the client, while Unsloth loads and serves models via a \*\*local API\*\*.After setup, OpenClaw will run against your local model through Unsloth, letting you use it directly as an \*\*AI agent.\*\* In this tutorial, we'll use \[Qwen3.6\](/docs/models/qwen3.6.md). [Connecting to OpenClaw](https://unsloth.ai/pages/CwQEpEmkKPmyEYdnEngt#connecting-to-openclaw) [Quickstart](https://unsloth.ai/pages/CwQEpEmkKPmyEYdnEngt#quickstart) {% hint style="info" %} In this tutorial, we’ll use \`unsloth/Qwen3.6-27B-GGUF\` in Unsloth and access it through OpenClaw. Prefer a different model? Swap in any other model by loading it in Unsloth and updating the configuration. {% endhint %} ### Installing OpenClaw {% tabs %} {% tab title="macOS, Linux, WSL" %} Install OpenClaw using the official installer: \`curl -fsSL https://openclaw.ai/install.sh | bash\` This sets up OpenClaw and guides you through initial setup. {% endtab %} {% tab title="Windows (PowerShell)" %} Install OpenClaw using the official installer: \`iwr -useb https://openclaw.ai/install.ps1 | iex\` This sets up OpenClaw and guides you through initial setup. {% endtab %} {% endtabs %} ### ⚡ Quickstart After installing OpenClaw, we'll need to install Unsloth Studio to enable OpenClaw to serve and run inference of local models. 1. \*\*Install or update\*\* \[\*\*Unsloth Studio\*\*\](/docs/new/studio.md)\*\*.\*\* Earlier versions don't expose the external API. See Installation. 2. \*\*Launch Unsloth.\*\* Note the port it starts on is usually \`8000\` or \`8888\`. You'll see it in the terminal output and in the browser URL (\`http://localhost:PORT\`). 3. \*\*Load a model.\*\* Click \*\*New Chat\*\*, pick or search a model (GGUF), and wait for it to finish loading. 4. \*\*Connect OpenClaw.\*\* Run \`unsloth start openclaw\` to launch OpenClaw with the loaded Unsloth model in a separate managed environment. Your normal OpenClaw configuration is left unchanged. ### ⚙️ Launch OpenClaw with \`unsloth start\` OpenClaw can connect to a model already running in Unsloth Studio, or start one automatically when Unsloth is not running. Connect to a running Unsloth instance Once a model is loaded in Unsloth Studio, run: \`\`\`bash unsloth start openclaw \`\`\` ![](https://unsloth.ai/files/neb4bX02lAk8QyGOtGCR) Unsloth launches OpenClaw in local TUI mode with the Unsloth provider, model, and context length configured inside a separate managed environment. Your normal OpenClaw setup is left unchanged. Once OpenClaw opens, give it a task such as: \`\`\` inspect this repo and summarize it in one sentence: https://github.com/unslothai/unsloth \`\`\` ![](https://unsloth.ai/files/6LaBkYeUqTFXxa2QogLz) \#### Start a model automatically If Unsloth Studio is not already running, pass a model ID: \`\`\`bash unsloth start openclaw --model unsloth/Qwen3.6-27B-GGUF \`\`\` Unsloth starts the model, launches OpenClaw, and stops the temporary server when you exit. \\ \\ To keep your OpenClaw state and return to the same session later, add \`--persist\` and choose a session key: \`\`\`bash unsloth start openclaw \\ --model unsloth/Qwen3.6-27B-GGUF \\ --persist \\ agent --local \\ --session-key my-session \\ --message "Inspect this repository" \`\`\` Continue later using the same session key: \`\`\`bash unsloth start openclaw \\ --model unsloth/Qwen3.6-27B-GGUF \\ --persist \\ agent --local \\ --session-key my-session \\ --message "Continue" \`\`\` See the complete \[unsloth start\](/docs/integrations/unsloth-start.md) reference for named sessions, model loading, remote Unsloth servers, and advanced options. The rest of this guide covers the manual OpenClaw provider setup. ### Manual OpenClaw provider setup ### 🔑 Creating an API key Keys are created from \*\*Unsloth → Settings → API Keys\*\*. 1. Open the sidebar, click your \*\*Unsloth\*\* avatar at the bottom-left. 2. Go to \*\*Settings\*\* → \*\*API Keys\*\*. 3. Enter a friendly name (e.g. \`claude-code-macbook\`). 4. \*(Optional)\* Set an expiry. 5. Click \*\*Create\*\*. 6. \*\*Copy the key immediately.\*\* Unsloth stores only a hash and you won't be able to view it again. ![](https://unsloth.ai/files/mIewhCcJSWNVe9g92qw6) All keys start with the \`sk-unsloth-\` prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with \`401 Unauthorized\`. ### Connecting to OpenClaw OpenClaw reads its config from \`~/.openclaw/openclaw.json\`. Add (or merge) a \`models\` block with a \`unsloth\` provider pointing at Unsloth's Anthropic Messages API. ![](https://unsloth.ai/files/eGHdaU73Dw3RPGRj07KM) {% code title="\\~/.openclaw/openclaw\\.json" %} \`\`\`json { "models": { "mode": "merge", "providers": { "unsloth": { "baseUrl": "http://localhost:8888", "apiKey": "sk-unsloth-xxxxxxxxxxxx", "api": "anthropic-messages", "models": \[ { "id": "unsloth/Qwen3.6-27B-GGUF", "name": "unsloth/Qwen3.6-27B-GGUF" } \], "authHeader": true } } } } \`\`\` {% endcode %} \*\*Notes:\*\* \* \`baseUrl\` is the Unsloth origin with no path. OpenClaw talks to Unsloth over the Anthropic Messages API, and the Anthropic SDK appends \`/v1/messages\` itself, so do not add \`/v1\` here (a trailing \`/v1\` would send requests to \`/v1/v1/messages\`). \* \`api: "anthropic-messages"\` tells OpenClaw to talk to Unsloth's \`/v1/messages\` endpoint. \* \`authHeader: true\` sends your key as \`Authorization: Bearer …\`. \* Set each model's \`id\` and \`name\` to the name you chose when loading the model in Unsloth. \* If you're running Unsloth on a remote machine, replace \`localhost:8888\` with that machine's address (e.g. \`http://10.0.0.42:8888\`). ### Optional: configure model behavior OpenClaw connects through the model running in Unsloth. Runtime settings can be configured when starting the server. \`\`\`bash # Configure default generation behavior (--disable-tools passes OpenClaw's own tools through) unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF \\ --disable-tools \\ --reasoning off \\ --temp 0.6 \`\`\` {% hint style="warning" %} Use \`--disable-tools\` when driving OpenClaw (or any external coding agent). By default Unsloth Studio runs its own server-side tools, which swallows the agent's tool calls, so OpenClaw answers but never edits files. \`--disable-tools\` switches to passthrough, so OpenClaw's own tools are used. {% endhint %} Use \`--reasoning off\` to turn thinking off, or \`--reasoning on\` to turn it on for models that support reasoning. \`\`\`bash # Allow connections from other devices unsloth run \\ --model unsloth/gemma-4-26B-A4B-it-GGUF \\ -H 0.0.0.0 \\ -p 8888 \`\`\` This starts the server on \`0.0.0.0:8888\`, allowing other devices on your local network to connect. For more advanced runtime configuration, see the main \[API tuning\](https://unsloth.ai/docs/basics/api#unsloth-run-command) section. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/openclaw.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/gspo-reinforcement-learning.md). # GSPO Reinforcement Learning We're introducing GSPO which is a variant of \[GRPO\](/docs/get-started/reinforcement-learning-rl-guide.md#from-rlhf-ppo-to-grpo-and-rlvr) made by the Qwen team at Alibaba. They noticed the observation that when GRPO takes importance weights for each token, even though inherently advantages do not scale or change with each token. This lead to the creation of GSPO, which now assigns the importance on the sequence likelihood rather than the individual token likelihoods of the tokens. \* Use our free GSPO notebooks for: \[\*\*gpt-oss-20b\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) and \[\*\*Qwen2.5-VL\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen2\_5\_7B\_VL\_GRPO.ipynb) Enable GSPO in Unsloth by setting \`importance\_sampling\_level = "sequence"\` in the GRPO config. The difference between these two algorithms can be seen below, both from the GSPO paper from Qwen and Alibaba: ![](https://unsloth.ai/files/7ntPBNOCiL5uh506pLzt) GRPO Algorithm, Source: [Qwen](https://arxiv.org/abs/2507.18071) ![](https://unsloth.ai/files/eyxi56SIssHb6e4xuA7L) GSPO algorithm, Source: [Qwen](https://arxiv.org/abs/2507.18071) In Equation 1, it can be seen that the advantages scale each of the rows into the token logprobs before that tensor is sumed. Essentially, each token is given the same scaling even though that scaling was given to the entire sequence rather than each individual token. A simple diagram of this can be seen below: ![](https://unsloth.ai/files/ebkGBYfegbkcvTt89vou) GRPO Logprob Ratio row wise scaled with advantages Equation 2 shows that the logprob ratios for each sequence is summed and exponentiated after the Logprob ratios are computed, and only the resulting now sequence ratios get row wise multiplied by the advantages. ![](https://unsloth.ai/files/jPCxEjDztfNea4N75aF3) GSPO Sequence Ratio row wise scaled with advantages Enabling GSPO is simple, all you need to do is set the \`importance\_sampling\_level = "sequence"\` flag in the GRPO config. \`\`\`python training\_args = GRPOConfig( output\_dir = "vlm-grpo-unsloth", per\_device\_train\_batch\_size = 8, gradient\_accumulation\_steps = 4, learning\_rate = 5e-6, adam\_beta1 = 0.9, adam\_beta2 = 0.99, weight\_decay = 0.1, warmup\_ratio = 0.1, lr\_scheduler\_type = "cosine", optim = "adamw\_8bit", # beta = 0.00, epsilon = 3e-4, epsilon\_high = 4e-4, num\_generations = 8, max\_prompt\_length = 1024, max\_completion\_length = 1024, log\_completions = False, max\_grad\_norm = 0.1, temperature = 0.9, # report\_to = "none", # Set to "wandb" if you want to log to Weights & Biases num\_train\_epochs = 2, # For a quick test run, increase for full training # GSPO is below: importance\_sampling\_level = "sequence", # Dr GRPO / GAPO etc loss\_type = "dr\_grpo", ) \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/gspo-reinforcement-learning.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl.md). # FP16 vs BF16 for RL ### Float16 vs Bfloat16 There was a paper titled "\*\*Defeating the Training-Inference Mismatch via FP16\*\*" showing how using float16 precision can dramatically be better than using bfloat16 when doing reinforcement learning. ![](https://unsloth.ai/files/MjMadQKEiFioxkhHzebI) In fact the longer the generation, the worse it gets when using bfloat16: ![](https://unsloth.ai/files/5KjSqtAXMkjGu1sUZkd0) We did an investigation, and \*\*DO find float16 to be more stable\*\* than bfloat16 with much smaller gradient norms see and {% columns %} {% column width="50%" %} ![](https://unsloth.ai/files/30eDH3r96TWcsa6vWjS5) ![](https://unsloth.ai/files/734ILjvD3NTYnWhKZQmm) {% endcolumn %} {% column width="50%" %} ![](https://unsloth.ai/files/UU8dAMDG4zWQ9ox2p4rJ) {% endcolumn %} {% endcolumns %} ### :exploding\\\_head:A100 Cascade Attention Bug As per and , older vLLM versions (before 0.11.0) had broken attention mechanisms for A100 and similar GPUs. Please update vLLM! We also by default disable cascade attention in vLLM during Unsloth reinforcement learning if we detect an older vLLM version. ![](https://unsloth.ai/files/OmroYSGwJl7jFVmc5TTR) Different hardware also changes results, where newer and more expensive GPUs have less KL difference between the inference and training sides: ![](https://unsloth.ai/files/ag8GrDaJLuqAt6PAqRh9) \### :fire:Using float16 in Unsloth RL To use float16 precision in Unsloth GRPO and RL, you just need to set \`dtype = torch.float16\` and we'll take care of the rest! {% code overflow="wrap" %} \`\`\`python from unsloth import FastLanguageModel import torch max\_seq\_length = 2048 # Can increase for longer reasoning traces lora\_rank = 32 # Larger rank = smarter, but slower model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/Qwen3-4B-Base", max\_seq\_length = max\_seq\_length, load\_in\_4bit = False, # False for LoRA 16bit fast\_inference = True, # Enable vLLM fast inference max\_lora\_rank = lora\_rank, gpu\_memory\_utilization = 0.9, # Reduce if out of memory dtype = torch.float16, # Use torch.float16, torch.bfloat16 ) \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/training-ai-agents-with-rl.md). # Training AI Agents with RL “Agentic” AI is becoming more popular over time. In this context, an “agent” is an LLM that is given a high-level goal and a set of tools to achieve it. Agents are also typically “multi-turn” — they can perform an action, see what effect it had on the environment, and then perform another action repeatedly, until they achieve their goal or fail trying. Unfortunately, even very capable LLMs can have a hard time performing complex multi-turn agentic tasks reliably. Interestingly, we’ve found that training agents using an RL algorithm called \[GRPO (Group Relative Policy Optimization)\](/docs/get-started/reinforcement-learning-rl-guide/tutorial-train-your-own-reasoning-model-with-grpo.md) can make them far more reliable! In this guide, you will learn how to to build reliable AI agents using open-source tools. ## 🎨 Training RL Agents with ART \[ART (Agent Reinforcement Trainer)\](https://github.com/openpipe/art) built on top of \[Unsloth\](https://github.com/unslothai/unsloth)’s GRPOTrainer, is a tool that makes training multi-turn agents possible and easy. If you’re already using Unsloth for GRPO and need to train agents that can handle complex, multi-turn interactions, ART simplifies the process. ![](https://unsloth.ai/files/bvI4a0WwZDSBW6zpBgdI) Agent models trained with Unsloth+ART are often able to outperform prompted models on agentic workflows. \### ART + Unsloth ART builds on top of Unsloth’s memory- and compute-efficient GRPO implementation. In addition, it adds the following capabilities: #### 1. Multi-Turn Agent Training ART introduces the concept of a “trajectory”, which is built up as your agent executes. These trajectories can then be scored and used for GRPO. Trajectories can be complex, and even include non-linear histories, sub-agent calls, etc. They also support tool calls and responses. #### 2. Flexible Integration into Existing Codebases If you already have an agent working with a prompted model, ART tries to minimize the number of changes you need to make to wrap your existing agent loop and use it for training. Architecturally, ART is split into a “frontend” client that lives in your codebase and communicates via API with a “backend” where the actual training happens (these can also be colocated on a single machine if you prefer using ART’s \`LocalBackend\`). This gives some key benefits: \* \*\*Minimal setup required\*\*: The ART frontend is has minimal dependencies and can be easily added to existing Python codebases. \* \*\*Train from anywhere\*\*: You can run the ART client on your laptop and let the ART server kick off an ephemeral GPU-enabled environment, or run on a local GPU \* \*\*OpenAI-compatible API\*\*: The ART backend serves your model undergoing training via an OpenAI-compatible API, which is compatible with most existing codebases. #### 3. RULER: Zero-Shot Agent Rewards ART also provides a built-in general-purpose reward function called \[RULER\](https://art.openpipe.ai/fundamentals/ruler) (Relative Universal LLM-Elicited Rewards), which can eliminate the need for hand-crafted reward functions. Surprisingly, agents RL-trained with the RULER automatic reward function often match or surpass the performance of agents trained using hand-written reward functions. This makes getting started with RL easier. ![](https://unsloth.ai/files/uoHUIxuLtTFcUhpVDtuc) \`\`\`python # Before: Hours of reward engineering def complex\_reward\_function(trajectory): # 50+ lines of careful scoring logic... pass # After: One line with RULER judged\_group = await ruler\_score\_group(group, "openai/o3") \`\`\` ### When to Choose ART ART might be a good fit for projects that need: 1. \*\*Multi-step agent capabilities\*\*: When your use case involves agents that need to take multiple actions, use tools, or have extended conversations 2. \*\*Rapid prototyping without reward engineering\*\*: RULER’s automatic reward scoring can cut your project’s development time by 2-3x 3. \*\*Integration with existing systems\*\*: When you need to add RL capabilities to an existing agentic codebase with minimal changes ### Code Example: ART in Action \`\`\`python import art from art.rewards import ruler\_score\_group # Initialize model with Unsloth-supported basemodel model = art.TrainableModel( name="agent-001", project="my-agentic-task", base\_model="Qwen/Qwen2.5-14B-Instruct", # Any Unsloth-supported model ) # Define your rollout function async def rollout(model: art.Model, scenario: Scenario) -> art.Trajectory: openai\_client = model.openai\_client() trajectory = art.Trajectory( messages\_and\_choices=\[ {"role": "system", "content": "..."}, {"role": "user", "content": "..."} \] ) # Your agent logic here... return trajectory # Train with RULER for automatic rewards groups = await art.gather\_trajectory\_groups( ( art.TrajectoryGroup(rollout(model, scenario) for \_ in range(8)) for scenario in scenarios ), after\_each=lambda group: ruler\_score\_group( group, "openai/o3", swallow\_exceptions=True ) ) await model.train(groups) \`\`\` ### Getting Started To add ART to your Unsloth-based project: \`\`\`bash pip install openpipe-art # or \`uv add openpipe-art\` \`\`\` Then check out the \[example notebooks\](https://art.openpipe.ai/getting-started/notebooks) to see ART in action with tasks like: \* Email retrieval agents that beat o3 \* Game-playing agents (2048, Tic Tac Toe, Codenames) \* Complex reasoning tasks (Temporal Clue) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/training-ai-agents-with-rl.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/dgx-station.md). # Fine-Tuning LLMs on NVIDIA DGX Station with Unsloth You can now train LLMs locally on your NVIDIA DGX Station with \[Unsloth\](https://github.com/unslothai/unsloth). DGX Station has more than \*\*\\~200GB VRAM\*\* and over \*\*700GB of unified GPU / CPU memory\*\* and combines a Grace CPU and a Blackwell GPU in a tightly connected system designed for large-scale AI workloads. Linked by NVLink-C2C, the CPU and GPU remain distinct but work together far more efficiently than in a traditional CPU-GPU setup. In this guide, we’ll use Unsloth notebooks train \[Qwen3.5\](#qwen3.5-35b-a3b-fine-tuning) and \[gpt-oss-120b\](#gpt-oss-120b-fine-tuning) on DGX Station. Thank you to NVIDIA for providing some early access DGX Station hardware to test Unsloth on! ### Quickstart You will need \`python3\` installed and in particular the dev headers are needed. On our system we have \`python 3.12\` so we will install the 3.12 dev headers. \`\`\`bash sudo apt update sudo apt install python3.12-dev \`\`\` Then create a fresh virtual environment to install \[Unsloth\](https://github.com/unslothai/unsloth). This way we minimize dependency conflicts and preserve the state of the current working environment. {% code overflow="wrap" %} \`\`\`bash python3 -m venv .unsloth source .unsloth/bin/activate pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 \`\`\` {% endcode %} {% hint style="warning" %} First install \`torch\` from the \`cuda 13\` index otherwise we could get the CPU version or a mismatch in architecture and capabilities! {% endhint %} ![](https://unsloth.ai/files/E9hIEDEOYjKrxOhYGfoK) ![](https://unsloth.ai/files/xHFjS2LspAMat3aPITuD) Now we can install Unsloth: \`\`\`bash pip install unsloth \`\`\` ![](https://unsloth.ai/files/hwoftFeLeRbNxBARTKUv) ![](https://unsloth.ai/files/Ujfe3OBKTGapf4EMFgFI) Now lets install \`xformers\` and optionally build \`flash-attention\` from source. Both packages take time so please be patient while they build. {% code overflow="wrap" expandable="true" %} \`\`\`bash pip install --no-deps --no-build-isolation xformers==0.0.33.post1 # Optionally flash-attn # Clone and build (targets sm\_100 for B300) git clone https://github.com/Dao-AILab/flash-attention cd flash-attention # B300 = sm\_100, set arch explicitly TORCH\_CUDA\_ARCH\_LIST="10.0" MAX\_JOBS=8 pip install . --no-build-isolation cd .. \`\`\` {% endcode %} ![](https://unsloth.ai/files/fJwWtucuRuGV5IOwTIl8) ![](https://unsloth.ai/files/wVIlLnxqkLowUzwULOt8) {% columns %} {% column %} For Qwen 3.5 MoE we’ll want to download two kernel packages \`flash-linear-attention\` and \`causal-conv1d\` to make it fast. {% code overflow="wrap" expandable="true" %} \`\`\`bash pip install --no-build-isolation flash-linear-attention causal\_conv1d==1.6.0 \`\`\` {% endcode %} {% endcolumn %} {% column %} ![](https://unsloth.ai/files/s8P5hHdw1fv1lMae1EYq) {% endcolumn %} {% endcolumns %} If you don’t already have a notebook client, install one. For this guide we will use Jupyter Notebook: {% code overflow="wrap" expandable="true" %} \`\`\`bash cd .. pip install notebook pip install ipywidgets \`\`\` {% endcode %} Finally we download the actual Unsloth notebooks to run. There are 250+ notebooks for LLM Training as well as Python scripts. {% code overflow="wrap" expandable="true" %} \`\`\`bash git clone https://github.com/unslothai/notebooks.git cd notebooks \`\`\` {% endcode %} ### Training Tutorials {% columns %} {% column %} Now we can launch Jupyter Notebook and navigate to the UI on a browser. {% code overflow="wrap" expandable="true" %} \`\`\`bash jupyter notebook \`\`\` {% endcode %} {% endcolumn %} {% column %} ![](https://unsloth.ai/files/LFMstEbEucYUN0YXRzKF) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} Copy and paste the \`localhost\` site with token parameter and paste into your browser. You should see something like: The \`nb\` folder has all the notebooks to run. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/R0WmyNvDb4AMOiHnGSO2) {% endcolumn %} {% endcolumns %} #### Qwen3.5-35B-A3B Training {% columns %} {% column %} Open the file \`nb/Qwen3\_5\_MoE.ipynb\`. Skip past the installation section since we already installed everything we need beforehand. Navigate to the Unsloth section and start executing cells from there. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/4TUFsk2w1oNkFObIBqma) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} The notebook covers model setup, dataset preparation, and trainer configuration. Each step can take some time as we are downloading a very large model, initializing billions of weights, and further optimizing to make it run fast. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/RJPjKDkOnQPiwkunc9ec) {% endcolumn %} {% endcolumns %} Training is very fast with the default setting. On the DGX Station there is plenty of memory so you can play with the default training hyper parameters to really push the memory and compute. Once done training you can save the model for later, push the model to Hugging Face Hub to share with others, or export to a quantized format. #### gpt-oss-120b Training {% columns %} {% column %} Open the file \`nb/gpt-oss-(120B)\_A100-Fine-tuning.ipynb\`. Skip past the installation section since we already installed the prerequisites and navigate to the Unsloth section. We can start running the notebook from there. The notebook will use around 72 GB of GPU memory and take about 10 minutes. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/j6TWAjyTH0VIZq7fPdUf) {% endcolumn %} {% endcolumns %} {% columns %} {% column %} Each cell can take some time to run as we need to download the model, initialize the weights, and further optimize for a fast experience. The notebook goes through dataset preprocessing and trainer setup. Once we get to the \`trainer.train()\` cell and execute training begins. {% endcolumn %} {% column %} ![](https://unsloth.ai/files/LCGRG6uOJmeIggRDmtE4) {% endcolumn %} {% endcolumns %} Now that it’s complete we can save the model for later use, push to Hugging Face Hub to share with the world, or export it to GGUF format. ![](https://unsloth.ai/files/6fNfHR8ZKBnNMMs76c4i) Read more about NVIDIA's DGX Station at \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/dgx-station.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/vllm-engine-arguments.md). # vLLM Engine Arguments vLLM engine arguments, flags, options for serving models on vLLM. | Argument | Example and use-case | | --- | --- | | **`--gpu-memory-utilization`** | Default 0.9. How much VRAM usage vLLM can use. Reduce if going out of memory. Try setting this to 0.95 or 0.97. | | **`--max-model-len`** | Set maximum sequence length. Reduce this if going out of memory! For example set **`--max-model-len 32768`** to use only 32K sequence lengths. | | **`--quantization`** | Use fp8 for dynamic float8 quantization. Use this in tandem with **`--kv-cache-dtype`** fp8 to enable float8 KV cache as well. | | **`--kv-cache-dtype`** | Use `fp8` for float8 KV cache to reduce memory usage by 50%. | | **`--port`** | Default is 8000. How to access vLLM's localhost ie http://localhost:8000 | | **`--api-key`** | Optional - Set the password (or no password) to access the model. | | **`--tensor-parallel-size`** | Default is 1. Splits model across tensors. Set this to how many GPUs you are using - if you have 4, set this to 4. 8, then 8. You should have NCCL, otherwise this might be slow. | | **`--pipeline-parallel-size`** | Default is 1. Splits model across layers. Use this with **`--pipeline-parallel-size`** where TP is used within each node, and PP is used across multi-node setups (set PP to number of nodes) | | **`--enable-lora`** | Enables LoRA serving. Useful for serving Unsloth finetuned LoRAs. | | **`--max-loras`** | How many LoRAs you want to serve at 1 time. Set this to 1 for 1 LoRA, or say 16. This is a queue so LoRAs can be hot-swapped. | | **`--max-lora-rank`** | Maximum rank of all LoRAs. Possible choices are `8`, `16`, `32`, `64`, `128`, `256`, `320`, `512` | | **`--dtype`** | Allows `auto`, `bfloat16`, `float16` Float8 and other quantizations use a different flag - see `--quantization` | | **`--tokenizer`** | Specify the tokenizer path like `unsloth/gpt-oss-20b` if the served model has a different tokenizer. | | **`--hf-token`** | Add your HuggingFace token if needed for gated models | | **`--swap-space`** | Default is 4GB. CPU offloading usage. Reduce if you have VRAM, or increase for low memory GPUs. | | **`--seed`** | Default is 0 for vLLM | | **`--disable-log-stats`** | Disables logging like throughput, server requests. | | **`--enforce-eager`** | Disables compilation. Faster to load, but slower for inference. | | **`--disable-cascade-attn`** | Useful for Reinforcement Learning runs for vLLM < 0.11.0, as Cascade Attention was slightly buggy on A100 GPUs (Unsloth fixes this) | \### :tada:Float8 Quantization For example to host Llama 3.3 70B Instruct (supports 128K context length) with Float8 KV Cache and quantization, try: \`\`\`bash vllm serve unsloth/Llama-3.3-70B-Instruct \\ --quantization fp8 \\ --kv-cache-dtype fp8 \\ --gpu-memory-utilization 0.97 \\ --max-model-len 65536 \`\`\` ### :shaved\\\_ice:LoRA Hot Swapping / Dynamic LoRAs To enable LoRA serving for at most 4 LoRAs at 1 time (these are hot swapped / changed), first set the environment flag to allow hot swapping: See our \[LoRA Hot Swapping Guide\](/docs/basics/inference-and-deployment/vllm-guide/lora-hot-swapping-guide.md) for more details. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/vllm-engine-arguments.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/chat-templates.md). # Chat Templates In our GitHub, we have a list of every chat template Unsloth uses including for Llama, Mistral, Phi-4 etc. So if you need any pointers on the formatting or use case, you can view them here: \[github.com/unslothai/unsloth/blob/main/unsloth/chat\\\_templates.py\](https://github.com/unslothai/unsloth/blob/main/unsloth/chat\_templates.py) #### List of Colab chat template notebooks: \* \[Conversational\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* \[ChatML\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3\_\\(8B\\)-Ollama.ipynb) \* \[Ollama\](https://colab.research.google.com/drive/1WZDi7APtQ9VsvOrQSSC5DDtxq159j8iZ?usp=sharing) \* \[Text Classification\](https://github.com/timothelaborie/text\_classification\_scripts/blob/main/unsloth\_classification.ipynb) by Timotheeee \* \[Multiple Datasets\](https://colab.research.google.com/drive/1njCCbE1YVal9xC83hjdo2hiGItpY\_D6t?usp=sharing) by Flail ### Adding new tokens Unsloth has a function called \`add\_new\_tokens\` which allows you to add new tokens to your finetune. For example if you want to add \`\`, \`\` and \`\` we can do the following: \`\`\`python model, tokenizer = FastLanguageModel.from\_pretrained(...) from unsloth import add\_new\_tokens add\_new\_tokens(model, tokenizer, new\_tokens = \["", "", ""\]) model = FastLanguageModel.get\_peft\_model(...) \`\`\` {% hint style="warning" %} Note - you MUST always call \`add\_new\_tokens\` before \`FastLanguageModel.get\_peft\_model\`! {% endhint %} ## Multi turn conversations An issue if you didn't notice is the Alpaca dataset is single turn, whilst remember using ChatGPT was interactive and you can talk to it in multiple turns. For example, the left is what we want, but the right which is the Alpaca dataset only provides singular conversations. We want the finetuned language model to somehow learn how to do multi turn conversations just like ChatGPT. ![](https://unsloth.ai/files/e2H8w0IaXjcnsLUaMlGw) So we introduced the \`conversation\_extension\` parameter, which essentially selects some random rows in your single turn dataset, and merges them into 1 conversation! For example, if you set it to 3, we randomly select 3 rows and merge them into 1! Setting them too long can make training slower, but could make your chatbot and final finetune much better! ![](https://unsloth.ai/files/9i1v0jyc7wQuCeIrDfS0) Then set \`output\_column\_name\` to the prediction / output column. For the Alpaca dataset, it would be the output column. We then use the \`standardize\_sharegpt\` function to just make the dataset in a correct format for finetuning! Always call this! ![](https://unsloth.ai/files/11KZXsIRTu34NUa3rNGU) \## Customizable Chat Templates We can now specify the chat template for finetuning itself. The very famous Alpaca format is below: ![](https://unsloth.ai/files/zoaars71bDZlmuZDAcEN) But remember we said this was a bad idea because ChatGPT style finetunes require only 1 prompt? Since we successfully merged all dataset columns into 1 using Unsloth, we essentially can create the below style chat template with 1 input column (instruction) and 1 output: ![](https://unsloth.ai/files/nayJRk7BI85fnfr4IpPa) We just require you must put a \`{INPUT}\` field for the instruction and an \`{OUTPUT}\` field for the model's output field. We in fact allow an optional \`{SYSTEM}\` field as well which is useful to customize a system prompt just like in ChatGPT. For example, below are some cool examples which you can customize the chat template to be: ![](https://unsloth.ai/files/JIyHlsc6u41kEcYeCmZX) For the ChatML format used in OpenAI models: ![](https://unsloth.ai/files/jiws9K56jcnJS8eWWVix) Or you can use the Llama-3 template itself (which only functions by using the instruct version of Llama-3): We in fact allow an optional \`{SYSTEM}\` field as well which is useful to customize a system prompt just like in ChatGPT. ![](https://unsloth.ai/files/n2H5qEgArQG1GK8mJz4T) Or in the Titanic prediction task where you had to predict if a passenger died or survived in this Colab notebook which includes CSV and Excel uploading: ![](https://unsloth.ai/files/1qjcUI0HOUTFc4TA20je) \## Applying Chat Templates with Unsloth For datasets that usually follow the common chatml format, the process of preparing the dataset for training or finetuning, consists of four simple steps: \* Check the chat templates that Unsloth currently supports:\\\\ \`\`\` from unsloth.chat\_templates import CHAT\_TEMPLATES print(list(CHAT\_TEMPLATES.keys())) \`\`\` \\ This will print out the list of templates currently supported by Unsloth. Here is an example output:\\\\ \`\`\` \['unsloth', 'zephyr', 'chatml', 'mistral', 'llama', 'vicuna', 'vicuna\_old', 'vicuna old', 'alpaca', 'gemma', 'gemma\_chatml', 'gemma2', 'gemma2\_chatml', 'llama-3', 'llama3', 'phi-3', 'phi-35', 'phi-3.5', 'llama-3.1', 'llama-31', 'llama-3.2', 'llama-3.3', 'llama-32', 'llama-33', 'qwen-2.5', 'qwen-25', 'qwen25', 'qwen2.5', 'phi-4', 'gemma-3', 'gemma3'\] \`\`\` \\\\ \* Use \`get\_chat\_template\` to apply the right chat template to your tokenizer:\\\\ \`\`\` from unsloth.chat\_templates import get\_chat\_template tokenizer = get\_chat\_template( tokenizer, chat\_template = "gemma-3", # change this to the right chat\_template name ) \`\`\` \\\\ \* Define your formatting function. Here's an example:\\\\ \`\`\` def formatting\_prompts\_func(examples): convos = examples\["conversations"\] texts = \[tokenizer.apply\_chat\_template(convo, tokenize = False, add\_generation\_prompt = False) for convo in convos\] return { "text" : texts, } \`\`\` \\ \\ This function loops through your dataset applying the chat template you defined to each sample.\\\\ \* Finally, let's load the dataset and apply the required modifications to our dataset: \\\\ \`\`\` # Import and load dataset from datasets import load\_dataset dataset = load\_dataset("repo\_name/dataset\_name", split = "train") # Apply the formatting function to your dataset using the map method dataset = dataset.map(formatting\_prompts\_func, batched = True,) \`\`\` \\ If your dataset uses the ShareGPT format with "from"/"value" keys instead of the ChatML "role"/"content" format, you can use the \`standardize\_sharegpt\` function to convert it first. The revised code will now look as follows:\\ \\\\ \`\`\` # Import dataset from datasets import load\_dataset dataset = load\_dataset("mlabonne/FineTome-100k", split = "train") # Convert your dataset to the "role"/"content" format if necessary from unsloth.chat\_templates import standardize\_sharegpt dataset = standardize\_sharegpt(dataset) # Apply the formatting function to your dataset using the map method dataset = dataset.map(formatting\_prompts\_func, batched = True,) \`\`\` ## More Information Assuming your dataset is a list of list of dictionaries like the below: \`\`\`python \[ \[{'from': 'human', 'value': 'Hi there!'}, {'from': 'gpt', 'value': 'Hi how can I help?'}, {'from': 'human', 'value': 'What is 2+2?'}\], \[{'from': 'human', 'value': 'What's your name?'}, {'from': 'gpt', 'value': 'I'm Daniel!'}, {'from': 'human', 'value': 'Ok! Nice!'}, {'from': 'gpt', 'value': 'What can I do for you?'}, {'from': 'human', 'value': 'Oh nothing :)'},\], \] \`\`\` You can use our \`get\_chat\_template\` to format it. Select \`chat\_template\` to be any of \`zephyr, chatml, mistral, llama, alpaca, vicuna, vicuna\_old, unsloth\`, and use \`mapping\` to map the dictionary values \`from\`, \`value\` etc. \`map\_eos\_token\` allows you to map \`<|im\_end|>\` to EOS without any training. \`\`\`python from unsloth.chat\_templates import get\_chat\_template tokenizer = get\_chat\_template( tokenizer, chat\_template = "chatml", # Supports zephyr, chatml, mistral, llama, alpaca, vicuna, vicuna\_old, unsloth mapping = {"role" : "from", "content" : "value", "user" : "human", "assistant" : "gpt"}, # ShareGPT style map\_eos\_token = True, # Maps <|im\_end|> to instead ) def formatting\_prompts\_func(examples): convos = examples\["conversations"\] texts = \[tokenizer.apply\_chat\_template(convo, tokenize = False, add\_generation\_prompt = False) for convo in convos\] return { "text" : texts, } pass from datasets import load\_dataset dataset = load\_dataset("philschmid/guanaco-sharegpt-style", split = "train") dataset = dataset.map(formatting\_prompts\_func, batched = True,) \`\`\` You can also make your own custom chat templates! For example our internal chat template we use is below. You must pass in a \`tuple\` of \`(custom\_template, eos\_token)\` where the \`eos\_token\` must be used inside the template. \`\`\`python unsloth\_template = \\ "{{ bos\_token }}"\\ "{{ 'You are a helpful assistant to the user\\n' }}"\\ ""\\ " "\\ " "\\ "{{ '>>> User: ' + message\['content'\] + '\\n' }}"\\ " "\\ "{{ '>>> Assistant: ' + message\['content'\] + eos\_token + '\\n' }}"\\ " "\\ " "\\ " "\\ "{{ '>>> Assistant: ' }}"\\ " " unsloth\_eos\_token = "eos\_token" tokenizer = get\_chat\_template( tokenizer, chat\_template = (unsloth\_template, unsloth\_eos\_token,), # You must provide a template and EOS token mapping = {"role" : "from", "content" : "value", "user" : "human", "assistant" : "gpt"}, # ShareGPT style map\_eos\_token = True, # Maps <|im\_end|> to instead ) \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/chat-templates.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/lora-hot-swapping-guide.md). # LoRA Hot Swapping Guide ### :shaved\\\_ice: vLLM LoRA Hot Swapping / Dynamic LoRAs To enable LoRA serving for at most 4 LoRAs at 1 time (these are hot swapped / changed), first set the environment flag to allow hot swapping: \`\`\`bash export VLLM\_ALLOW\_RUNTIME\_LORA\_UPDATING=True \`\`\` Then, serve it with LoRA support: \`\`\`bash export VLLM\_ALLOW\_RUNTIME\_LORA\_UPDATING=True vllm serve unsloth/Llama-3.1-8B-Instruct \\ --quantization fp8 \\ --kv-cache-dtype fp8 \\ --gpu-memory-utilization 0.8 \\ --max-model-len 65536 \\ --enable-lora \\ --max-loras 4 \\ --max-lora-rank 64 \`\`\` To load a LoRA dynamically (set the lora name as well), do: \`\`\`bash curl -X POST http://localhost:8000/v1/load\_lora\_adapter \\ -H "Content-Type: application/json" \\ -d '{ "lora\_name": "LORA\_NAME", "lora\_path": "/path/to/LORA" }' \`\`\` To remove it from the pool: \`\`\`bash curl -X POST http://localhost:8000/v1/unload\_lora\_adapter \\ -H "Content-Type: application/json" \\ -d '{ "lora\_name": "LORA\_NAME" }' \`\`\` For example when finetuning with Unsloth: {% code overflow="wrap" %} \`\`\`python from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/Llama-3.1-8B-Instruct", max\_seq\_length = 2048, load\_in\_4bit = True, ) model = FastLanguageModel.get\_peft\_model(model) \`\`\` {% endcode %} Then after training, we save the LoRAs: \`\`\`python model.save\_pretrained("finetuned\_lora") tokenizer.save\_pretrained("finetuned\_lora") \`\`\` We can then load the LoRA: {% code overflow="wrap" %} \`\`\`bash curl -X POST http://localhost:8000/v1/load\_lora\_adapter \\ -H "Content-Type: application/json" \\ -d '{ "lora\_name": "LORA\_NAME\_finetuned\_lora", "lora\_path": "finetuned\_lora" }' \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/vllm-guide/lora-hot-swapping-guide.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/openai.md). # Connect OpenAI to Unsloth: Run GPT Models in Local Chat Learn how to connect OpenAI models, including GPT-5.5 to \[Unsloth\](https://github.com/unslothai/unsloth) so you can chat with all of them in an open-source local UI chat interface. By connecting your OpenAI API key, you can run GPT models inside Unsloth with features like \[web search\](#web-search-and-thinking), tool-calling, \[code execution\](#code-execution), \[image generation\](#image-generation), reusable code containers, and \[prompt caching\](#prompt-caching). This guide walks you through creating an OpenAI API key, connecting OpenAI as a provider, loading available models, and troubleshooting common setup issues. ### Setup {% stepper %} {% step %} #### Create an OpenAI API key Create an API key from the \[OpenAI dashboard\](https://platform.openai.com/api-keys). ![](https://unsloth.ai/files/PXaINS0IoRIME98oPVwT) {% endstep %} {% step %} #### Configure Connections Next, connect your provider to Unsloth. 1. Open \*\*Settings\*\* → \*\*Connections\*\*, then click \*\*Add Connection.\*\* 2. Select the OpenAI, then paste the API key you copied earlier. 3. Click \*\*Reload Models\*\* to refresh the list with models available to your account. 4. Choose the models you want to enable, then hit save. ![](https://unsloth.ai/files/wyW65uOZY4CSsSkvex2G) {% endstep %} {% step %} #### Ready to Chat The models you enabled will now appear under Connected in the Select Model dropdown. Supported GPT models can expose extra controls including image generation, thinking, web search and code execution. ![](https://unsloth.ai/files/Q6wnsLueufwOLm7uhaEM) {% endstep %} {% endstepper %} ### Code Execution When enabled, supported OpenAI models can run code in a provider sandbox to solve problems, analyse data, and work with files. ![](https://unsloth.ai/files/V90WGRM0SPKT5K6QUEQj) {% columns %} {% column width="50%" %} OpenAI uses reusable shell containers. In \*\*Code Execution\*\* settings, you can set the idle timeout, create containers, select the active container, refresh the list, or delete old containers. Select the same container in a new thread to continue with its files and state. {% endcolumn %} {% column width="50%" %} ![](https://unsloth.ai/files/uQSaDCf9MocZjwTlbJ2u) {% endcolumn %} {% endcolumns %} ### Prompt Caching {% columns %} {% column width="66.66666666666666%" %} Prompt caching reduces latency and cost when requests reuse the same long prefix. It is supported for compatible providers and servers, including OpenAI. Use the \*\*Prompt caching\*\* setting in the side panel to control caching behaviour for supported connections. {% endcolumn %} {% column width="33.33333333333334%" %} ![](https://unsloth.ai/files/QSJQ82qfp3z5H35zIou6) {% endcolumn %} {% endcolumns %} ![](https://unsloth.ai/files/rd7uqDkUz6YnddW01aRl) \### Web Search & Thinking Provider-side web search is available for supported models from OpenAI. The \*\*Think\*\* control adapts to the selected model: some models use an on/off toggle, while reasoning-effort models use model specific thinking levels. ![](https://unsloth.ai/files/46johOqXvwqWOhGnbso4) \### Image Generation Just like GPT, Unsloth also supports image generation. You can directly edit an image by clicking the “Edit Image” button and entering a new prompt to refine or regenerate it. Images are generated automatically when requested, but you can toggle this behavior off. A download button is also available, allowing you to save the image in its original full resolution. ![](https://unsloth.ai/files/tc7WCuUGdy9DA5PQwvDs) ![](https://unsloth.ai/files/GgH0XVMZeBOPsDM3rQVv) \### Troubleshooting If OpenAI fails to connect, check that the API key is valid and belongs to the correct OpenAI account. If a model does not appear after clicking \*\*Load Models\*\*, it may not be available for your account. You can enter the model ID manually or choose another model. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/openai.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf/speculative-decoding.md). # Speculative Decoding ## :llama:Speculative Decoding in llama.cpp, llama-server Speculative decoding in llama.cpp can be easily enabled via \`llama-cli\` and \`llama-server\` via the \`--model-draft\` argument. Note you must have a draft model, which generally is a smaller model, but it must have the same tokenizer ### Spec Decoding for GLM 4.7 \`\`\`python # !pip install huggingface\_hub hf\_transfer import os os.environ\["HF\_HUB\_ENABLE\_HF\_TRANSFER"\] = "0" # Can sometimes rate limit, so set to 0 to disable from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/GLM-4.7-GGUF", local\_dir = "unsloth/GLM-4.7-GGUF", allow\_patterns = \["\*UD-Q2\_K\_XL\*"\], # Dynamic 2bit Use "\*UD-TQ1\_0\*" for Dynamic 1bit ) snapshot\_download( repo\_id = "unsloth/GLM-4.5-Air-GGUF", local\_dir = "unsloth/GLM-4.5-Air-GGUF", allow\_patterns = \["\*UD-Q4\_K\_XL\*"\], # Dynamic 4bit. Use "\*UD-TQ1\_0\*" for Dynamic 1bit ) \`\`\` {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/GLM-4.7-GGUF/UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ --threads -1 \\ --fit on \\ --prio 3 \\ --temp 1.0 \\ --top-p 0.95 \\ --ctx-size 16384 \\ --jinja \`\`\` {% endcode %} ![](https://unsloth.ai/files/XfxG6Yzs08yZYdVXCIpE) {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\ --model unsloth/GLM-4.7-GGUF/UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ --model-draft unsloth/GLM-4.5-Air-GGUF/UD-Q4\_K\_XL/GLM-4.5-Air-UD-Q4\_K\_XL-00001-of-00002.gguf \\ --threads -1 \\ --fit on \\ --prio 3 \\ --temp 1.0 \\ --top-p 0.95 \\ --ctx-size 16384 \\ --ctx-size-draft 16384 \\ --jinja \\ --device CUDA0 \\ --device-draft CUDA0,CUDA1 \`\`\` {% endcode %} {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-server \\ --model unsloth/GLM-4.7-GGUF/UD-Q2\_K\_XL/GLM-4.7-UD-Q2\_K\_XL-00001-of-00003.gguf \\ --alias "unsloth/GLM-4.7" \\ --threads -1 \\ --fit on \\ --prio 3 \\ --temp 1.0 \\ --top-p 0.95 \\ --ctx-size 16384 \\ --port 8001 \\ --jinja \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf/speculative-decoding.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-ollama.md). # Saving models to Ollama See our guide below for the complete process on how to save models to \[Ollama\](https://github.com/ollama/ollama): {% content-ref url="/pages/cECKVbf1TpF5j7WC0riJ" %} \[Tutorial: Finetune Llama-3 and Use In Ollama\](/docs/get-started/fine-tuning-llms-guide/tutorial-how-to-finetune-llama-3-and-use-in-ollama.md) {% endcontent-ref %} ### Saving on Google Colab You can save the finetuned model as a small 100MB file called a LoRA adapter like below. You can instead push to the Hugging Face hub as well if you want to upload your model! Remember to get a Hugging Face token via: and add your token! ![](https://unsloth.ai/files/ckdSoKSGiZHxuYsmRLGX) After saving the model, we can again use Unsloth to run the model itself! Use \`FastLanguageModel\` again to call it for inference! ![](https://unsloth.ai/files/9r0QCOEc1oEapNdThVDc) \### Exporting to Ollama Finally we can export our finetuned model to Ollama itself! First we have to install Ollama in the Colab notebook: ![](https://unsloth.ai/files/IaeU19nq1f6sPoBLxgPV) Then we export the finetuned model we have to llama.cpp's GGUF formats like below: ![](https://unsloth.ai/files/lYNuzY8dAjmnVSyDXQPu) Reminder to convert \`False\` to \`True\` for 1 row, and not change every row to \`True\`, or else you'll be waiting for a very time! We normally suggest the first row getting set to \`True\`, so we can export the finetuned model quickly to \`Q8\_0\` format (8 bit quantization). We also allow you to export to a whole list of quantization methods as well, with a popular one being \`q4\_k\_m\`. Head over to to learn more about GGUF. We also have some manual instructions of how to export to GGUF if you want here: You will see a long list of text like below - please wait 5 to 10 minutes!! ![](https://unsloth.ai/files/aY5S6oEY9p5fQxGC0z5N) And finally at the very end, it'll look like below: ![](https://unsloth.ai/files/bPLVzW7MTJ9OAYsouNpV) Then, we have to run Ollama itself in the background. We use \`subprocess\` because Colab doesn't like asynchronous calls, but normally one just runs \`ollama serve\` in the terminal / command prompt. ![](https://unsloth.ai/files/mOVQVVOon5xdcm0fMaRW) \### Automatic \`Modelfile\` creation The trick Unsloth provides is we automatically create a \`Modelfile\` which Ollama requires! This is a just a list of settings and includes the chat template which we used for the finetune process! You can also print the \`Modelfile\` generated like below: ![](https://unsloth.ai/files/vxoOMjrFc32mTR6SwBrI) We then ask Ollama to create a model which is Ollama compatible, by using the \`Modelfile\` ![](https://unsloth.ai/files/ylhNcN12696Gz4Sf0Fba) \### Ollama Inference And we can now call the model for inference if you want to do call the Ollama server itself which is running on your own local machine / in the free Colab notebook in the background. Remember you can edit the yellow underlined part. ![](https://unsloth.ai/files/oaqz3wZ2zMVyaLT8q7K8) \### Running in Unsloth works well, but after exporting & running on Ollama, the results are poor You might sometimes encounter an issue where your model runs and produces good results on Unsloth, but when you use it on another platform like Ollama, the results are poor or you might get gibberish, endless/infinite generations \*or\* repeated outputs\*\*.\*\* \* The most common cause of this error is using an \*\*incorrect chat template\*\*\*\*.\*\* It’s essential to use the SAME chat template that was used when training the model in Unsloth and later when you run it in another framework, such as llama.cpp or Ollama. When inferencing from a saved model, it's crucial to apply the correct template. \* You must use the correct \`eos token\`. If not, you might get gibberish on longer generations. \* It might also be because your inference engine adds an unnecessary "start of sequence" token (or the lack of thereof on the contrary) so ensure you check both hypotheses! \* \*\*Use our conversational notebooks to force the chat template - this will fix most issues.\*\* \* Qwen-3 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* Gemma-3 4B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) \* Llama-3.2 3B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* Phi-4 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* Mistral v0.3 7B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \*\*More notebooks in our\*\* \[\*\*notebooks docs\*\*\](/docs/get-started/unsloth-notebooks.md) --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-ollama.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md). # Fine-tuning LLMs with Blackwell, RTX 50 series & Unsloth Unsloth now supports NVIDIA’s Blackwell architecture GPUs, including RTX 50-series GPUs (5060–5090), RTX PRO 6000, and GPUS such as B200, B40, GB100, GB102 and more! You can read the official \[NVIDIA blogpost here\](https://developer.nvidia.com/blog/train-an-llm-on-an-nvidia-blackwell-desktop-with-unsloth-and-scale-it/). Unsloth is now compatible with every NVIDIA GPU from 2018+ including the \[DGX Spark\](/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth.md). > \*\*Our new\*\* \[\*\*Docker image\*\*\](#docker) \*\*supports Blackwell. Run the Docker image and start training!\*\* \[\*\*Guide\*\*\](/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md) ### Pip install Simply install Unsloth: \`\`\`bash pip install unsloth \`\`\` If you see issues, another option is to create a separate isolated environment: \`\`\`bash python -m venv unsloth source unsloth/bin/activate pip install unsloth \`\`\` Note it might be \`pip3\` or \`pip3.13\` and also \`python3\` or \`python3.13\` You might encounter some Xformers issues, in which cause you should build from source: {% code overflow="wrap" %} \`\`\`bash # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH\_CUDA\_ARCH\_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. \`\`\` {% endcode %} ### Docker \[\*\*\`unsloth/unsloth\`\*\*\](https://hub.docker.com/r/unsloth/unsloth) is Unsloth's only Docker image. For Blackwell and 50-series GPUs, use this same image - no separate image needed. For installation instructions, please follow our \[Unsloth Docker guide\](/docs/blog/how-to-fine-tune-llms-with-unsloth-and-docker.md). ### uv \`\`\`bash uv pip install unsloth \`\`\` #### uv (Advanced) The installation order is important, since we want the overwrite bundled dependencies with specific versions (namely, \`xformers\` and \`triton\`). 1. I prefer to use \`uv\` over \`pip\` as it's faster and better for resolving dependencies, especially for libraries which depend on \`torch\` but for which a specific \`CUDA\` version is required per this scenario. Install \`uv\` \`\`\`bash curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env \`\`\` Create a project dir and venv: \`\`\`bash mkdir 'unsloth-blackwell' && cd 'unsloth-blackwell' uv venv .venv --python=3.12 --seed source .venv/bin/activate \`\`\` 2. Install \`vllm\` \`\`\`bash uv pip install -U vllm --torch-backend=cu128 \`\`\` Note that we have to specify \`cu128\`, otherwise \`vllm\` will install \`torch==2.7.0\` but with \`cu126\`. 3. Install \`unsloth\` dependencies \`\`\`bash uv pip install unsloth unsloth\_zoo bitsandbytes \`\`\` If you notice weird resolving issues due to Xformers, you can also install Unsloth from source without Xformers: \`\`\`bash uv pip install -qqq \\ "unsloth\_zoo\[base\] @ git+https://github.com/unslothai/unsloth-zoo" \\ "unsloth\[base\] @ git+https://github.com/unslothai/unsloth" \`\`\` 4. Download and build \`xformers\` (Optional) Xformers is optional, but it is definitely faster and uses less memory. We'll use PyTorch's native SDPA if you do not want Xformers. Building Xformers from source might be slow, so beware! \`\`\`bash # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH\_CUDA\_ARCH\_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. \`\`\` Note that we have to explicitly set \`TORCH\_CUDA\_ARCH\_LIST=12.0\`. 5. \`transformers\` Install any transformers version, but best to get the latest. \`\`\`bash uv pip install -U transformers \`\`\` ### Conda or mamba (Advanced) 1. Install \`conda/mamba\` \`\`\`bash curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh" \`\`\` Run the installation script \`\`\`bash bash Miniforge3-$(uname)-$(uname -m).sh \`\`\` Create a conda or mamba environment \`\`\`bash conda create --name unsloth-blackwell python==3.12 -y \`\`\` Activate newly created environment \`\`\`bash conda activate unsloth-blackwell \`\`\` 2. Install \`vllm\` Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your \`(unsloth-blackwell)user@machine:\` \`\`\`bash pip install -U vllm --extra-index-url https://download.pytorch.org/whl/cu128 \`\`\` Note that we have to specify \`cu128\`, otherwise \`vllm\` will install \`torch==2.7.0\` but with \`cu126\`. 3. Install \`unsloth\` dependencies Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your \`(unsloth-blackwell)user@machine:\` \`\`\`bash pip install unsloth unsloth\_zoo bitsandbytes \`\`\` 4. Download and build \`xformers\` (Optional) Xformers is optional, but it is definitely faster and uses less memory. We'll use PyTorch's native SDPA if you do not want Xformers. Building Xformers from source might be slow, so beware! You should see the name of your environment as a prefix to your terminal shell like this your \`(unsloth-blackwell)user@machine:\` \`\`\`bash # First uninstall xformers installed by previous libraries pip uninstall xformers -y # Clone and build pip install ninja export TORCH\_CUDA\_ARCH\_LIST="12.0" git clone --depth=1 https://github.com/facebookresearch/xformers --recursive cd xformers && python setup.py install && cd .. \`\`\` Note that we have to explicitly set \`TORCH\_CUDA\_ARCH\_LIST=12.0\`. 5. Update \`triton\` Make sure you are inside the activated conda/mamba environment. You should see the name of your environment as a prefix to your terminal shell like this your \`(unsloth-blackwell)user@machine:\` \`\`\`bash pip install -U "triton>=3.3.1" \`\`\` \`triton>=3.3.1\` is required for \`Blackwell\` support. 6. \`Transformers\` Install any transformers version, but best to get the latest. \`\`\`bash uv pip install -U transformers \`\`\` If you are using mamba as your package just replace conda with mamba for all commands shown above. ### WSL-Specific Notes If you're using WSL (Windows Subsystem for Linux) and encounter issues during xformers compilation (reminder Xformers is optional, but faster for training) follow these additional steps: 1. \*\*Increase WSL Memory Limit\*\* Create or edit the WSL configuration file: \`\`\`bash # Create or edit .wslconfig in your Windows user directory # (typically C:\\Users\\YourUsername\\.wslconfig) # Add these lines to the file \[wsl2\] memory=16GB # Minimum 16GB recommended for xformers compilation processors=4 # Adjust based on your CPU cores swap=2GB localhostForwarding=true \`\`\` After making these changes, restart WSL: \`\`\`powershell wsl --shutdown \`\`\` 2. \*\*Install xformers\*\* Use the following command to install xformers with optimized compilation for WSL: \`\`\`bash # Set CUDA architecture for Blackwell GPUs export TORCH\_CUDA\_ARCH\_LIST="12.0" # Install xformers from source with optimized build flags pip install -v --no-build-isolation -U git+https://github.com/facebookresearch/xformers.git@main#egg=xformers \`\`\` The \`--no-build-isolation\` flag helps avoid potential build issues in WSL environments. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/ollama.md). # How to Connect Ollama to Unsloth Ollama lets you run local LLMs on your own hardware, and \[Unsloth\](https://github.com/unslothai/unsloth) makes it easy to connect and run those models directly into a open-source UI chat interface. In this guide, you’ll learn how to install Ollama, run native Ollama models or GGUF models from Hugging Face, connect Ollama to Unsloth, and start chatting with local AI models. Whether you want to use models like \[Qwen\](/docs/models/qwen3.6.md), import a GGUF file, or expose your local Ollama server through an OpenAI-compatible endpoint, this walkthrough covers the full setup from installation to first chat. ### Setup {% stepper %} {% step %} #### Install or prepare Ollama {% tabs %} {% tab title="macOS" %} Install Ollama with the install script: \`\`\`bash curl -fsSL https://ollama.com/install.sh | sh \`\`\` You can also download Ollama manually from \[ollama.com/download\](https://ollama.com/download). {% endtab %} {% tab title="Windows" %} Install Ollama from PowerShell: \`\`\`powershell irm https://ollama.com/install.ps1 | iex \`\`\` You can also download Ollama manually from \[ollama.com/download\](https://ollama.com/download/OllamaSetup.exe). {% endtab %} {% tab title="Linux" %} Install Ollama with the install script: \`\`\`bash curl -fsSL https://ollama.com/install.sh | sh \`\`\` You can also download Ollama manually from \[ollama.com/download\](https://docs.ollama.com/linux#manual-install). {% endtab %} {% tab title="Docker" %} The official Ollama Docker image is \`ollama/ollama\` on Docker Hub. \`\`\`bash docker run -d \\ -v ollama:/root/.ollama \\ -p 11434:11434 \\ --name ollama \\ ollama/ollama \`\`\` {% endtab %} {% endtabs %} Ollama usually runs at: \`\`\` http://localhost:11434 \`\`\` {% endstep %} {% step %} #### Run a model You can choose a model in two common ways: \* Search native Ollama models at \[ollama.com/search\](https://ollama.com/search), then copy the model name. \* Use a GGUF model from Hugging Face, then copy the Ollama command from \*\*Use this model\*\*. For an Ollama model, pull and run it: \`\`\`bash ollama pull qwen3.6:35b-a3b ollama run qwen3.6:35b-a3b \`\`\` If the Ollama app or service is not already running, start it first: \`\`\`bash ollama serve \`\`\` #### Pick a GGUF from Hugging Face If you are using a GGUF model from Hugging Face, the easiest way to get the command is from the model page. Open the model you want to use, click \*\*Use this model\*\*, then choose \*\*Ollama\*\* from the local apps list. Pick the quantization you want from the dropdown, then copy the generated command. ![](https://unsloth.ai/files/fkqT2O0gZJUueFitKrDy) For example, with Ollama: \`\`\`bash ollama run hf.co/unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4\_K\_XL \`\`\` This helps avoid mistakes with the repo name or quantization tag. {% endstep %} {% step %} #### Connect Ollama to Unsloth Open \*\*Settings → Connections\*\*, then click \*\*Add Connection\*\*. Select \*\*Ollama\*\*, then enter your connection details: ![](https://unsloth.ai/files/dtryqjag6SQ8zkeOIp05) Use the Ollama URL shown in the Unsloth form. In most local setups, this is: \`\`\` http://localhost:11434 \`\`\` If Unsloth asks for an OpenAI-compatible base URL, use: \`\`\` http://localhost:11434/v1 \`\`\` Ollama normally does not need an API key. Leave the API key field empty unless you are using a proxy that requires one. Click \*\*Load Models\*\* to fetch the models running in Ollama, or enter the \*\*model ID\*\* yourself, for example \`qwen3.6\`. {% endstep %} {% step %} #### Ready to Chat After you click \*\*Add Connection\*\*, the models you enabled will now appear under \*\*Connected\*\* in the \*\*Select Model\*\* dropdown. {% endstep %} {% endstepper %} #### Common Ollama commands Use these while setting up the model you want to expose to Unsloth: | Command | What it does | | ----------------------------- | ---------------------------------------- | | \`ollama run qwen3.6:35b-a3b\` | Run a model and open an interactive chat | | \`ollama pull qwen3.6:35b-a3b\` | Download a model without starting chat | | \`ollama ls\` | List downloaded models | | \`ollama ps\` | List models currently running | | \`ollama stop qwen3.6:35b-a3b\` | Stop a running model | | \`ollama rm qwen3.6:35b-a3b\` | Remove a downloaded model | | \`ollama serve\` | Start the Ollama server | If you are importing a local GGUF into Ollama, create a \`Modelfile\`, then run: \`\`\`bash ollama create -f Modelfile \`\`\` If Ollama is not detected, make sure the Ollama app or service is running. Then click \*\*Load Models\*\* again in Unsloth. For the full command list, see the \[Ollama CLI reference\](https://docs.ollama.com/cli). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/ollama.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md). # Hugging Face Hub, XET debugging #### Downloads are stuck at 90% to 99% ![](https://unsloth.ai/files/6UOXjlldwlweDRQVkZ0q) If you see downloads via \`hf download unsloth/\*\` get stuck at 90% or 99% of progress for quite some time, cancel the current run, and try adding using below commands: \`\`\`bash pip install -U huggingface\_hub HF\_HOME=".cache\_new/huggingface" \\ HF\_XET\_CACHE=".cache\_new/huggingface/xet" \\ HF\_HUB\_CACHE=".cache\_new/huggingface/hub" \\ HF\_XET\_HIGH\_PERFORMANCE=1 \\ HF\_XET\_CHUNK\_CACHE\_SIZE\_BYTES=0 \\ HF\_XET\_RECONSTRUCT\_WRITE\_SEQUENTIALLY=0 \\ HF\_XET\_NUM\_CONCURRENT\_RANGE\_GETS=64 \\ hf download unsloth/Qwen3-Coder-Next-GGUF \\ --local-dir unsloth/Qwen3-Coder-Next-GGUF \\ --include "\*UD-Q6\_K\_XL\*" \`\`\` #### Rate limited or 429 Too Many Requests? Try using \`snapshot\_download\` instead, then import Unsloth which will set the correct Hugging Face variables for you: \`\`\`python import unsloth import os os.environ\["HF\_HOME"\] = ".cache\_new/huggingface" os.environ\["HF\_XET\_CACHE"\] = ".cache\_new/huggingface/xet" os.environ\["HF\_HUB\_CACHE"\] = ".cache\_new/huggingface/hub" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Coder-Next-GGUF", local\_dir = "unsloth/Qwen3-Coder-Next-GGUF", allow\_patterns = \["\*UD-Q6\_K\_XL\*"\], ) \`\`\` Or maybe try getting a Hugging Face token first via \`\`\`bash pip install -U huggingface\_hub HF\_HOME=".cache\_new/huggingface" \\ HF\_XET\_CACHE=".cache\_new/huggingface/xet" \\ HF\_HUB\_CACHE=".cache\_new/huggingface/hub" \\ HF\_XET\_HIGH\_PERFORMANCE=1 \\ HF\_XET\_CHUNK\_CACHE\_SIZE\_BYTES=0 \\ HF\_XET\_RECONSTRUCT\_WRITE\_SEQUENTIALLY=0 \\ HF\_XET\_NUM\_CONCURRENT\_RANGE\_GETS=64 \\ hf download unsloth/Qwen3-Coder-Next-GGUF \\ --local-dir unsloth/Qwen3-Coder-Next-GGUF \\ --include "\*UD-Q6\_K\_XL\*" \\ --token "hf\_ADD\_YOUR\_HUGGING\_FACE\_TOKEN\_HERE" \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners.md). # Fine-tuning for Beginners If you're a beginner, here might be the first questions you'll ask before your first fine-tune. You can also always ask our community by joining our \[Reddit page\](https://www.reddit.com/r/unsloth/). | | | | | | --- | --- | --- | --- | | [/pages/BAeSP6aOxvSeDUzCgKOK](https://unsloth.ai/pages/BAeSP6aOxvSeDUzCgKOK) | See every model you can run / train with Unsloth. | How to run GGUFs or train LLMs? | [/pages/BAeSP6aOxvSeDUzCgKOK](https://unsloth.ai/pages/BAeSP6aOxvSeDUzCgKOK) | | [/pages/nw2c1elNySGBBav8WP9B](https://unsloth.ai/pages/nw2c1elNySGBBav8WP9B) | Step-by-step on how to fine-tune! | Learn the core basics of training. | [/pages/nw2c1elNySGBBav8WP9B](https://unsloth.ai/pages/nw2c1elNySGBBav8WP9B) | | [/pages/BSShKhLoFNlGWO5cN8VJ](https://unsloth.ai/pages/BSShKhLoFNlGWO5cN8VJ) | Instruct or Base Model? | How big should my dataset be? | [/pages/BSShKhLoFNlGWO5cN8VJ](https://unsloth.ai/pages/BSShKhLoFNlGWO5cN8VJ) | | [/pages/HP82bIzgldwxWk3OSzVy](https://unsloth.ai/pages/HP82bIzgldwxWk3OSzVy) | What can fine-tuning do for me? | RAG vs. Fine-tuning? | [/pages/HP82bIzgldwxWk3OSzVy](https://unsloth.ai/pages/HP82bIzgldwxWk3OSzVy) | | [/pages/WbSfE0ITQYsNqERZwnbZ](https://unsloth.ai/pages/WbSfE0ITQYsNqERZwnbZ) | How do I install Unsloth locally? | How to update Unsloth? | [/pages/WbSfE0ITQYsNqERZwnbZ](https://unsloth.ai/pages/WbSfE0ITQYsNqERZwnbZ) | | [/pages/XgcpRfamZHmHnRHBnVE4](https://unsloth.ai/pages/XgcpRfamZHmHnRHBnVE4) | How do I structure/prepare my dataset? | How do I collect data? | | | [/pages/odJXZM9jv284RKqZ2Pna](https://unsloth.ai/pages/odJXZM9jv284RKqZ2Pna) | Does Unsloth work on my GPU? | How much VRAM will I need? | [/pages/odJXZM9jv284RKqZ2Pna](https://unsloth.ai/pages/odJXZM9jv284RKqZ2Pna) | | [/pages/gEugERiAw2ztDNt98JVR](https://unsloth.ai/pages/gEugERiAw2ztDNt98JVR) | How do I save my model locally? | How do I run my model via Ollama or vLLM? | [/pages/gEugERiAw2ztDNt98JVR](https://unsloth.ai/pages/gEugERiAw2ztDNt98JVR) | | [/pages/y6obKRSk8TwyjIrCjuGE](https://unsloth.ai/pages/y6obKRSk8TwyjIrCjuGE) | What happens when I change a parameter? | What parameters should I change? | | ![](https://unsloth.ai/files/4vqGG5CV8cphfk8Fag7E) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/fine-tuning-for-beginners.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth.md). # Multi-GPU Fine-tuning with Unsloth Unsloth currently supports multi-GPU setups through libraries like Accelerate and DeepSpeed. This means you can already leverage parallelism methods such as \*\*FSDP\*\* and \*\*DDP\*\* with Unsloth. #### \*\*See our new Distributed Data Parallel\*\* \[\*\*(DDP) multi-GPU Guide here\*\*\](/docs/basics/multi-gpu-training-with-unsloth/ddp.md)\*\*.\*\* We know that the process can be complex and requires manual setup. We’re working hard to make multi-GPU support much simpler and more user-friendly, and we’ll be announcing official multi-GPU support for Unsloth soon. For now, you can use our \[Magistral-2509 Kaggle notebook\](/docs/models/tutorials/magistral-how-to-run-and-fine-tune.md#fine-tuning-magistral-with-unsloth) as an example which utilizes multi-GPU Unsloth to fit the 24B parameter model or our \[DDP guide\](/docs/basics/multi-gpu-training-with-unsloth/ddp.md). \*\*In the meantime\*\*, to enable multi GPU for DDP, do the following: 1. Create your training script as \`train.py\` (or similar). For example, you can use one of our \[training scripts\](https://github.com/unslothai/notebooks/tree/main/python\_scripts) created from our various notebooks! 2. Run \`accelerate launch train.py\` or \`torchrun --nproc\_per\_node N\_GPUS train.py\` where \`N\_GPUS\` is the number of GPUs you have. #### \*\*Pipeline / model splitting loading\*\* If you do not have enough VRAM for 1 GPU to load say Llama 70B, no worries - we will split the model for you on each GPU! To enable this, use the \`device\_map = "balanced"\` flag: \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( "unsloth/Llama-3.3-70B-Instruct", load\_in\_4bit = True, device\_map = "balanced", ) \`\`\` \*\*Stay tuned for our official announcement!\*\*\\ For more details, check out our ongoing \[Pull Request\](https://github.com/unslothai/unsloth/issues/2435) discussing multi-GPU support. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/gpu-mode-conference.md). # GPU Mode - Reinforcement Learning Mini Conference 2026 Unsloth GitHub for efficient GRPO / GSPO with lots of custom kernel work: Here are our attached PDF slides: {% file src="/files/XP2Q9dkkXqKuDEMO1gdo" %} ### Reinforcement Learning Resources For step-by-step guides, beginner or advanced tutorials, you can refer to our docs for: {% columns %} {% column width="50%" %} {% content-ref url="/pages/vT6jTKG1LCfN7HoJ4fVR" %} \[Reinforcement Learning\](/docs/get-started/reinforcement-learning-rl-guide.md) {% endcontent-ref %} {% content-ref url="/pages/PVxTK4Eal3B77LmfhdRU" %} \[RL Reward Hacking\](/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/rl-reward-hacking.md) {% endcontent-ref %} {% endcolumn %} {% column width="50%" %} {% content-ref url="/pages/P9PfsQ0BjnZuwPaXA93E" %} \[Advanced RL Docs\](/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation.md) {% endcontent-ref %} {% content-ref url="/pages/wQboNZf1ZtBJ9Qk4WQxb" %} \[FP16 vs BF16 for RL\](/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/fp16-vs-bf16-for-rl.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} {% columns %} {% column width="50%" %} {% content-ref url="/pages/pFyRT83vVFiXANdxamCs" %} \[FP8 RL\](/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning.md) {% endcontent-ref %} {% endcolumn %} {% column width="50%" %} {% content-ref url="/pages/aV6S9cmDmSv5ky4ZCo5d" %} \[Vision RL\](/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} ### Live Video Recording {% embed url="" %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/gpu-mode-conference.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint.md). # Finetuning from Last Checkpoint You must edit the \`Trainer\` first to add \`save\_strategy\` and \`save\_steps\`. Below saves a checkpoint every 50 steps to the folder \`outputs\`. \`\`\`python trainer = SFTTrainer( .... args = TrainingArguments( .... output\_dir = "outputs", save\_strategy = "steps", save\_steps = 50, ), ) \`\`\` Then in the trainer do: \`\`\`python trainer\_stats = trainer.train(resume\_from\_checkpoint = True) \`\`\` Which will start from the latest checkpoint and continue training. ### Wandb Integration \`\`\` # Install library !pip install wandb --upgrade # Setting up Wandb !wandb login import os os.environ\["WANDB\_PROJECT"\] = "" os.environ\["WANDB\_LOG\_MODEL"\] = "checkpoint" \`\`\` Then in \`TrainingArguments()\` set \`\`\` report\_to = "wandb", logging\_steps = 1, # Change if needed save\_steps = 100 # Change if needed run\_name = "" # (Optional) \`\`\` To train the model, do \`trainer.train()\`; to resume training, do \`\`\` import wandb run = wandb.init() artifact = run.use\_artifact('//', type='model') artifact\_dir = artifact.download() trainer.train(resume\_from\_checkpoint=artifact\_dir) \`\`\` ## :question:How do I do Early Stopping? If you want to stop or pause the finetuning / training run since the evaluation loss is not decreasing, then you can use early stopping which stops the training process. Use \`EarlyStoppingCallback\`. As usual, set up your trainer and your evaluation dataset. The below is used to stop the training run if the \`eval\_loss\` (the evaluation loss) is not decreasing after 3 steps or so. \`\`\`python from trl import SFTConfig, SFTTrainer trainer = SFTTrainer( args = SFTConfig( fp16\_full\_eval = True, per\_device\_eval\_batch\_size = 2, eval\_accumulation\_steps = 4, output\_dir = "training\_checkpoints", # location of saved checkpoints for early stopping save\_strategy = "steps", # save model every N steps save\_steps = 10, # how many steps until we save the model save\_total\_limit = 3, # keep only 3 saved checkpoints to save disk space eval\_strategy = "steps", # evaluate every N steps eval\_steps = 10, # how many steps until we do evaluation load\_best\_model\_at\_end = True, # MUST USE for early stopping metric\_for\_best\_model = "eval\_loss", # metric we want to early stop on greater\_is\_better = False, # the lower the eval loss, the better ), model = model, tokenizer = tokenizer, train\_dataset = new\_dataset\["train"\], eval\_dataset = new\_dataset\["test"\], ) \`\`\` We then add the callback which can also be customized: \`\`\`python from transformers import EarlyStoppingCallback early\_stopping\_callback = EarlyStoppingCallback( early\_stopping\_patience = 3, # How many steps we will wait if the eval loss doesn't decrease # For example the loss might increase, but decrease after 3 steps early\_stopping\_threshold = 0.0, # Can set higher - sets how much loss should decrease by until # we consider early stopping. For eg 0.01 means if loss was # 0.02 then 0.01, we consider to early stop the run. ) trainer.add\_callback(early\_stopping\_callback) \`\`\` Then train the model as usual via \`trainer.train() .\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/finetuning-from-last-checkpoint.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/unsloth-benchmarks.md). # Unsloth Benchmarks \* For more detailed benchmarks, read our \[Llama 3.3 Blog\](https://unsloth.ai/blog/llama3-3). \* Benchmarking of Unsloth was also conducted by \[🤗Hugging Face\](https://huggingface.co/blog/unsloth-trl). {% hint style="warning" %} If your speed seems slower at first, it’s likely because \`torch.compile\` typically takes \\~5 minutes (or longer) to warm up and finish compiling. Make sure you measure throughput \*\*after\*\* it’s fully loaded as over longer runs, Unsloth should be much faster. {% endhint %} Tested on H100 and \[Blackwell\](/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth.md) GPUs. We tested using the Alpaca Dataset, a batch size of 2, gradient accumulation steps of 4, rank = 32, and applied QLoRA on all linear layers (q, k, v, o, gate, up, down): | Model | VRAM | 🦥Unsloth speed | 🦥VRAM reduction | 🦥Longer context | 😊Hugging Face + FA2 | | --- | --- | --- | --- | --- | --- | | Llama 3.3 (70B) | 80GB | 2x | \>75% | 13x longer | 1x | | Llama 3.1 (8B) | 80GB | 2x | \>70% | 12x longer | 1x | \## Context length benchmarks {% hint style="info" %} The more data you have, the less VRAM Unsloth uses due to our \[gradient checkpointing\](https://unsloth.ai/blog/long-context) algorithm + Apple's CCE algorithm! {% endhint %} ### \*\*Llama 3.1 (8B) max. context length\*\* We tested Llama 3.1 (8B) Instruct and did 4bit QLoRA on all linear layers (Q, K, V, O, gate, up and down) with rank = 32 with a batch size of 1. We padded all sequences to a certain maximum sequence length to mimic long context finetuning workloads. | GPU VRAM | 🦥Unsloth context length | Hugging Face + FA2 | | -------- | ------------------------ | ------------------ | | 8 GB | 2,972 | OOM | | 12 GB | 21,848 | 932 | | 16 GB | 40,724 | 2,551 | | 24 GB | 78,475 | 5,789 | | 40 GB | 153,977 | 12,264 | | 48 GB | 191,728 | 15,502 | | 80 GB | 342,733 | 28,454 | ### \*\*Llama 3.3 (70B) max. context length\*\* We tested Llama 3.3 (70B) Instruct on a 80GB A100 and did 4bit QLoRA on all linear layers (Q, K, V, O, gate, up and down) with rank = 32 with a batch size of 1. We padded all sequences to a certain maximum sequence length to mimic long context finetuning workloads. | GPU VRAM | 🦥Unsloth context length | Hugging Face + FA2 | | -------- | ------------------------ | ------------------ | | 48 GB | 12,106 | OOM | | 80 GB | 89,389 | 6,916 | --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/unsloth-benchmarks.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/conda-install.md). # Conda Install {% hint style="warning" %} Only use Conda if you have it. If not, use \[Pip\](/docs/get-started/install/pip-install.md). {% endhint %} Select either \`pytorch-cuda=11.8,12.1\` for CUDA 11.8 or CUDA 12.1. We support \`python=3.10,3.11,3.12\`. \`\`\`bash conda create --name unsloth\_env python=3.11 -y conda activate unsloth\_env pip install unsloth \`\`\` If you're looking to install Conda in a Linux environment, \[read here\](https://docs.anaconda.com/miniconda/), or run the below: {% code overflow="wrap" %} \`\`\`bash mkdir -p ~/miniconda3 wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86\_64.sh -O ~/miniconda3/miniconda.sh bash ~/miniconda3/miniconda.sh -b -u -p ~/miniconda3 rm -rf ~/miniconda3/miniconda.sh ~/miniconda3/bin/conda init bash ~/miniconda3/bin/conda init zsh \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/conda-install.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/unsloth-environment-flags.md). # Unsloth Environment Flags | Environment variable | Purpose | | | --- | --- | --- | | `os.environ["UNSLOTH_RETURN_LOGITS"] = "1"` | Forcibly returns logits - useful for evaluation if logits are needed. | | | `os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"` | Disables auto compiler. Could be useful to debug incorrect finetune results. | | | `os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"` | Disables fast generation for generic models. | | | `os.environ["UNSLOTH_ENABLE_LOGGING"] = "1"` | Enables auto compiler logging - useful to see which functions are compiled or not. | | | `os.environ["UNSLOTH_FORCE_FLOAT32"] = "1"` | On float16 machines, use float32 and not float16 mixed precision. Useful for Gemma 3. | | | `os.environ["UNSLOTH_STUDIO_DISABLED"] = "1"` | Disables extra features. | | | `os.environ["UNSLOTH_COMPILE_DEBUG"] = "1"` | Turns on extremely verbose `torch.compile`logs. | | | `os.environ["UNSLOTH_COMPILE_MAXIMUM"] = "0"` | Enables maximum `torch.compile`optimizations - not recommended. | | | `os.environ["UNSLOTH_COMPILE_IGNORE_ERRORS"] = "1"` | Can turn this off to enable fullgraph parsing. | | | `os.environ["UNSLOTH_FULLGRAPH"] = "0"` | Enable `torch.compile` fullgraph mode | | | `os.environ["UNSLOTH_DISABLE_AUTO_UPDATES"] = "1"` | Forces no updates to `unsloth-zoo` | | Another possibility is maybe the model uploads we uploaded are corrupted, but unlikely. Try the following: \`\`\`python model, tokenizer = FastVisionModel.from\_pretrained( "Qwen/Qwen2-VL-7B-Instruct", use\_exact\_model\_name = True, ) \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/unsloth-environment-flags.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/continued-pretraining.md). # Continued Pretraining \* The \[text completion notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_\\(7B\\)-Text\_Completion.ipynb) is for continued pretraining/raw text. \* The \[continued pretraining notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-CPT.ipynb) is for learning another language. You can read more about continued pretraining and our release in our \[blog post\](https://unsloth.ai/blog/contpretraining). ## What is Continued Pretraining? Continued or continual pretraining (CPT) is necessary to “steer” the language model to understand new domains of knowledge, or out of distribution domains. Base models like Llama-3 8b or Mistral 7b are first pretrained on gigantic datasets of trillions of tokens (Llama-3 for e.g. is 15 trillion). But sometimes these models have not been well trained on other languages, or text specific domains, like law, medicine or other areas. So continued pretraining (CPT) is necessary to make the language model learn new tokens or datasets. ## Advanced Features: ### Loading LoRA adapters for continued finetuning If you saved a LoRA adapter through Unsloth, you can also continue training using your LoRA weights. The optimizer state will be reset as well. To load even optimizer states to continue finetuning, see the next section. \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "LORA\_MODEL\_NAME", max\_seq\_length = max\_seq\_length, dtype = dtype, load\_in\_4bit = load\_in\_4bit, ) trainer = Trainer(...) trainer.train() \`\`\` ### Continued Pretraining & Finetuning the \`lm\_head\` and \`embed\_tokens\` matrices Add \`lm\_head\` and \`embed\_tokens\`. For Colab, sometimes you will go out of memory for Llama-3 8b. If so, just add \`lm\_head\`. \`\`\`python model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj", "lm\_head", "embed\_tokens",\], lora\_alpha = 16, ) \`\`\` Then use 2 different learning rates - a 2-10x smaller one for the \`lm\_head\` or \`embed\_tokens\` like so: \`\`\`python from unsloth import UnslothTrainer, UnslothTrainingArguments trainer = UnslothTrainer( .... args = UnslothTrainingArguments( .... learning\_rate = 5e-5, embedding\_learning\_rate = 5e-6, # 2-10x smaller than learning\_rate ), ) \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/continued-pretraining.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/unsloth-inference.md). # Unsloth Inference Unsloth supports natively 2x faster inference. For our inference only notebook, click \[here\](https://colab.research.google.com/drive/1aqlNQi7MMJbynFDyOQteD2t0yVfjb9Zh?usp=sharing). All QLoRA, LoRA and non LoRA inference paths are 2x faster. This requires no change of code or any new dependencies. from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name = "lora_model", # YOUR MODEL YOU USED FOR TRAINING max_seq_length = max_seq_length, dtype = dtype, load_in_4bit = load_in_4bit, ) FastLanguageModel.for_inference(model) # Enable native 2x faster inference text_streamer = TextStreamer(tokenizer) _ = model.generate(**inputs, streamer = text_streamer, max_new_tokens = 64) \#### NotImplementedError: A UTF-8 locale is required. Got ANSI Sometimes when you execute a cell \[this error\](https://github.com/googlecolab/colabtools/issues/3409) can appear. To solve this, in a new cell, run the below: \`\`\`python import locale locale.getpreferredencoding = lambda: "UTF-8" \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/unsloth-inference.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment.md). # Inference & Deployment You can also run your fine-tuned models by using \[Unsloth's 2x faster inference\](/docs/basics/inference-and-deployment/unsloth-inference.md). | | | | | --- | --- | --- | | [Unsloth Studio Chat](https://unsloth.ai/pages/qyazJc8QbOQ0mtlu6uEv#run-models-locally) | [/pages/FdMvLj95MbkAR4aHURvS](https://unsloth.ai/pages/FdMvLj95MbkAR4aHURvS) | | | [llama.cpp - Saving to GGUF](https://unsloth.ai/pages/T7ZPf3SNAwDykZNgXptE) | [/pages/T7ZPf3SNAwDykZNgXptE](https://unsloth.ai/pages/T7ZPf3SNAwDykZNgXptE) | [/pages/T7ZPf3SNAwDykZNgXptE](https://unsloth.ai/pages/T7ZPf3SNAwDykZNgXptE) | | [Unsloth API endpoint](https://unsloth.ai/pages/7sCtc6YWnJBYthTjQsr7) | [/pages/7sCtc6YWnJBYthTjQsr7](https://unsloth.ai/pages/7sCtc6YWnJBYthTjQsr7) | | | [vLLM](https://unsloth.ai/pages/fhJtaLFFXVsGnbMUiACo) | [/pages/fhJtaLFFXVsGnbMUiACo](https://unsloth.ai/pages/fhJtaLFFXVsGnbMUiACo) | [/pages/fhJtaLFFXVsGnbMUiACo](https://unsloth.ai/pages/fhJtaLFFXVsGnbMUiACo) | | [Ollama](https://unsloth.ai/pages/8UQUlu6UU8hhx3FiWc5B) | [/pages/8UQUlu6UU8hhx3FiWc5B](https://unsloth.ai/pages/8UQUlu6UU8hhx3FiWc5B) | [/pages/8UQUlu6UU8hhx3FiWc5B](https://unsloth.ai/pages/8UQUlu6UU8hhx3FiWc5B) | | [Connecting to a Provider](https://unsloth.ai/pages/4iT5ojqRctQslDUikk5N) | [/pages/4iT5ojqRctQslDUikk5N](https://unsloth.ai/pages/4iT5ojqRctQslDUikk5N) | | | [SGLang](https://unsloth.ai/pages/WehjgbuawqCXogREXvGG) | [/pages/WehjgbuawqCXogREXvGG](https://unsloth.ai/pages/WehjgbuawqCXogREXvGG) | [/pages/T8vAb3VMIaDyIUlVtrdK](https://unsloth.ai/pages/T8vAb3VMIaDyIUlVtrdK) | | [Troubleshooting](https://unsloth.ai/pages/e09iCJirEAJOrGDlyLre) | [/pages/e09iCJirEAJOrGDlyLre](https://unsloth.ai/pages/e09iCJirEAJOrGDlyLre) | [/pages/e09iCJirEAJOrGDlyLre](https://unsloth.ai/pages/e09iCJirEAJOrGDlyLre) | | [llama-server & OpenAI endpoint](https://unsloth.ai/pages/wqYCVI9bC4YRS7jlwWhR) | [/pages/wqYCVI9bC4YRS7jlwWhR](https://unsloth.ai/pages/wqYCVI9bC4YRS7jlwWhR) | | | [NVFP4](https://unsloth.ai/pages/d3MRkjfFC691sNPclO3q) | | | \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/amd/amd-hackathon.md). # AMD AI Reinforcement Learning Hackathon with Unsloth You can view Unsloth's GitHub repo here: Here is the link to our AMD fine-tuning notebooks: {% embed url="" %} {% code overflow="wrap" %} \`\`\`bash wget 'https://raw.githubusercontent.com/unslothai/notebooks/refs/heads/main/nb/gpt\_oss\_(20B)\_Reinforcement\_Learning\_2048\_Game\_BF16.ipynb' \`\`\` {% endcode %} If wanting to upgrade Unsloth / Unsloth Zoo: {% code overflow="wrap" %} \`\`\`bash uv pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm7.0 --upgrade --force-reinstall pip uninstall unsloth unsloth\_zoo -y && \\ pip install git+https://github.com/unslothai/unsloth-zoo git+https://github.com/unslothai/unsloth --no-deps --force-reinstall --no-cache-dir \`\`\` {% endcode %} For bitsandbytes: \`\`\`bash pip install "unsloth\[amd\] @ git+https://github.com/unslothai/unsloth" \`\`\` If you see: {% code overflow="wrap" %} \`\`\` error: Failed to install: bitsandbytes-1.33.7rc0-py3-none-manylinux\_2\_24\_x86\_64.whl (bitsandbytes==1.33.7rc0 (from https://github.com/bitsandbytes-foundation/bitsandbytes/releases/download/continuous-release\_main/bitsandbytes-1.33.7.preview-py3-none-manylinux\_2\_24\_x86\_64.whl)) Caused by: Wheel version does not match filename (0.49.2.dev0 != 1.33.7rc0), which indicates a malformed wheel. If this is intentional, set UV\_SKIP\_WHEEL\_FILENAME\_CHECK=1. \`\`\` {% endcode %} Do NOT use UV\\\_SKIP\\\_WHEEL\\\_FILENAME\\\_CHECK, instead ONLY use \`pip install "unsloth\[amd\] @ git+https://github.com/unslothai/unsloth"\` (NOT uv) since uv destroys bitsandbytes. Maybe add a check to the PRs if possible to catch these. For AMD installation instructions, you can view our guide here: {% content-ref url="/pages/GUxDmG8LaAiQinraCXCr" %} \[AMD\](/docs/get-started/install/amd.md) {% endcontent-ref %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/amd/amd-hackathon.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use.md). # What Model Should I Use for Fine-tuning? ## Llama, Qwen, Mistral, Phi or? When preparing for fine-tuning, one of the first decisions you'll face is selecting the right model. Here's a step-by-step guide to help you choose: {% stepper %} {% step %} \*\*Choose a model that aligns with your usecase\*\* \* E.g. For image-based training, select a vision model such as \*Llama 3.2 Vision\*. For code datasets, opt for a specialized model like \*Qwen Coder 2.5\*. \* \*\*Licensing and Requirements\*\*: Different models may have specific licensing terms and \[system requirements\](/docs/get-started/fine-tuning-for-beginners/unsloth-requirements.md#system-requirements). Be sure to review these carefully to avoid compatibility issues. {% endstep %} {% step %} \*\*Assess your storage, compute capacity and dataset\*\* \* Use our \[VRAM guideline\](/docs/get-started/fine-tuning-for-beginners/unsloth-requirements.md#approximate-vram-requirements-based-on-model-parameters) to determine the VRAM requirements for the model you’re considering. \* Your dataset will reflect the type of model you will use and amount of time it will take to train {% endstep %} {% step %} \*\*Select a Model and Parameters\*\* \* We recommend using the latest model for the best performance and capabilities. For instance, as of January 2025, the leading 70B model is \*Llama 3.3\*. \* You can stay up to date by exploring our \[model catalog\](/docs/get-started/unsloth-model-catalog.md) to find the newest and relevant options. {% endstep %} {% step %} \*\*Choose Between Base and Instruct Models\*\* Further details below: {% endstep %} {% endstepper %} ## Instruct or Base Model? When preparing for fine-tuning, one of the first decisions you'll face is whether to use an instruct model or a base model. ### Instruct Models Instruct models are pre-trained with built-in instructions, making them ready to use without any fine-tuning. These models, including GGUFs and others commonly available, are optimized for direct usage and respond effectively to prompts right out of the box. Instruct models work with conversational chat templates like ChatML or ShareGPT. ### \*\*Base Models\*\* Base models, on the other hand, are the original pre-trained versions without instruction fine-tuning. These are specifically designed for customization through fine-tuning, allowing you to adapt them to your unique needs. Base models are compatible with instruction-style templates like \[Alpaca or Vicuna\](/docs/basics/chat-templates.md), but they generally do not support conversational chat templates out of the box. ### Should I Choose Instruct or Base? The decision often depends on the quantity, quality, and type of your data: \* \*\*1,000+ Rows of Data\*\*: If you have a large dataset with over 1,000 rows, it's generally best to fine-tune the base model. \* \*\*300–1,000 Rows of High-Quality Data\*\*: With a medium-sized, high-quality dataset, fine-tuning the base or instruct model are both viable options. \* \*\*Less than 300 Rows\*\*: For smaller datasets, the instruct model is typically the better choice. Fine-tuning the instruct model enables it to align with specific needs while preserving its built-in instructional capabilities. This ensures it can follow general instructions without additional input unless you intend to significantly alter its functionality. \* For information how how big your dataset should be, \[see here\](/docs/get-started/fine-tuning-llms-guide/datasets-guide.md#how-big-should-my-dataset-be) ## Fine-tuning models with Unsloth You can change the model name to whichever model you like by matching it with model's name on Hugging Face e.g. 'unsloth/llama-3.1-8b-unsloth-bnb-4bit'. We recommend starting with \*\*Instruct models\*\*, as they allow direct fine-tuning using conversational chat templates (ChatML, ShareGPT etc.) and require less data compared to \*\*Base models\*\* (which uses Alpaca, Vicuna etc). Learn more about the differences between \[instruct and base models here\](#instruct-or-base-model). \* Model names ending in \*\*\`unsloth-bnb-4bit\`\*\* indicate they are \[\*\*Unsloth dynamic 4-bit\*\*\](https://unsloth.ai/blog/dynamic-4bit) \*\*quants\*\*. These models consume slightly more VRAM than standard BitsAndBytes 4-bit models but offer significantly higher accuracy. \* If a model name ends with just \*\*\`bnb-4bit\`\*\*, without "unsloth", it refers to a standard BitsAndBytes 4-bit quantization. \* Models with \*\*no suffix\*\* are in their original \*\*16-bit or 8-bit formats\*\*. While they are the original models from the official model creators, we sometimes include important fixes - such as chat template or tokenizer fixes. So it's recommended to use our versions when available. ### Experimentation is Key {% hint style="info" %} We recommend experimenting with both models when possible. Fine-tune each one and evaluate the outputs to see which aligns better with your goals. {% endhint %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/vs-code.md). # How to Fine-tune LLMs in VS Code with Unsloth & Colab GPUs You can now fine-tune LLMs directly from Visual Studio Code (VSCode), locally or by using Google Colab's extension. In this guide, you’ll learn how to use the open-source training \[repo: Unsloth\](https://github.com/unslothai/unsloth), to connect any \[fine-tuning notebook\](/docs/get-started/unsloth-notebooks.md) in VS Code to a Colab runtime, so you can train on your local or free Colab GPU. You can also view our video tutorial \[here\](#video-tutorial). {% stepper %} {% step %} ### VS Code and Colab Tutorial: To begin we will need to have: \* Installed \[VS Code\](https://code.visualstudio.com/). Git (for cloning the notebook repo) should be installed by default. \* A \*\*Google account\*\* (to authenticate with Colab) \* Recommended: \*\*Jupyter\*\* extension (most VS Code setups already have it) {% endstep %} {% step %} #### Install the Colab extension in VS Code 1. Open \*\*Extensions\*\* in VS Code (\`Ctrl+Shift+X\` / \`Cmd+Shift+X\`) 2. Search for \*\*“Colab”\*\* and install the \*\*Google Colab\*\* extension ![](https://unsloth.ai/files/ZtWzii7U5cueyybjcrJn) {% endstep %} {% step %} #### Open an Unsloth notebook 1. Clone the Unsloth \[notebooks repository\](https://github.com/unslothai/notebooks): \`\`\`bash git clone https://github.com/unslothai/notebooks cd notebooks/nb \`\`\` ![](https://unsloth.ai/files/k4UleubbtgEIByOCLIgr) 2\. Open your desired notebook. Unsloth supports most models including \[embedding\](/docs/basics/embedding-finetuning.md), \[TTS\](/docs/basics/text-to-speech-tts-fine-tuning.md). For example, we'll use Qwen3-4B RL: \`nb/Qwen3\_(4B)-GRPO.ipynb\` ![](https://unsloth.ai/files/Ed3GoENcIbE1yse0tngm) {% endstep %} {% step %} #### Select a kernel and choose Colab In the notebook toolbar, click \*\*Select Kernel\*\*, then choose \*\*Colab\*\* ![](https://unsloth.ai/files/6lYrGIPExWcPIRBlOz65) {% endstep %} {% step %} #### Add a new Colab server After choosing \*\*Colab\*\*, you’ll see a dropdown with server options. 1. Click \*\*+ Add New Colab Server\*\* 2. The first time, a browser window may open for Google authentication \* Log in, grant access, then return to VS Code ![](https://unsloth.ai/files/FO5UPj3fsM9VanRYURqW) {% endstep %} {% step %} #### Select GPU and name the server 1. Set \*\*Hardware accelerator\*\* to \*\*GPU\*\* 2. Choose a GPU type (for example \*\*T4\*\*, if available) 3. Give the server a name (anything you like) ![](https://unsloth.ai/files/cT283r3qcwvBBnUVtM0R) {% hint style="info" %} Note: GPU availability depends on your Colab plan and current capacity. If you don’t see GPU options, see troubleshooting below. {% endhint %} {% endstep %} {% step %} #### Pick the Python kernel Once connected to the Colab server, select the \*\*Python\*\* kernel that appears for that runtime (usually a Python 3 kernel). ![](https://unsloth.ai/files/hFH3mE73xnDE550DFjNN) {% endstep %} {% step %} #### Run the notebook \* Click \*\*Run All\*\* in the notebook toolbar (or run cells top-to-bottom) \* Watch the setup cells install dependencies and then start the Unsloth workflow \* You can view our dedicated \[fine-tuning\](/docs/get-started/fine-tuning-llms-guide.md) or \[reinforcement learning\](/docs/get-started/reinforcement-learning-rl-guide.md) guides for more info on exactly how to get started with Unsloth. {% endstep %} {% endstepper %} ### Video Tutorial {% embed url="" %} ### Troubleshooting #### After Colab server disconnects, the notebook won’t run on a new server \*\*What’s happening:\*\*\\ If the notebook stays open while the Colab server disconnects, VS Code can get stuck in a bad kernel/runtime state after reconnecting. Related \[GitHub issue\](https://github.com/googlecolab/colab-vscode/issues/200). \*\*Fix:\*\* Close the notebook tab completely and open the notebook again. #### You can’t select a GPU (only CPU shows up) Possible causes and fixes: \* \*\*Colab free tier capacity:\*\* GPUs may be temporarily unavailable → try again later. \* \*\*Not actually connected to a Colab runtime:\*\* re-check \*\*Select Kernel → Colab\*\* and ensure a Colab server is active. \* \*\*Account/region restrictions or limits reached:\*\* you may need to wait or use a different Google account / plan. #### Everything worked, but packages are “gone” after reconnecting Colab runtimes are \*\*ephemeral\*\*. When a server restarts, you usually need to re-run the setup/install cells (often the first few cells in the notebook). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/vs-code.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me.md). # FAQ + Is Fine-tuning Right For Me? ## Understanding Fine-Tuning Fine-tuning an LLM customizes its behavior, deepens its domain expertise, and optimizes its performance for specific tasks. By refining a pre-trained model (e.g. \*Llama-3.1-8B\*) with specialized data, you can: \* \*\*Update Knowledge\*\* – Introduce new, domain-specific information that the base model didn’t originally include. \* \*\*Customize Behavior\*\* – Adjust the model’s tone, personality, or response style to fit specific needs or a brand voice. \* \*\*Optimize for Tasks\*\* – Improve accuracy and relevance on particular tasks or queries your use-case requires. Think of fine-tuning as creating a specialized expert out of a generalist model. Some debate whether to use Retrieval-Augmented Generation (RAG) instead of fine-tuning, but fine-tuning can incorporate knowledge and behaviors directly into the model in ways RAG cannot. In practice, combining both approaches yields the best results - leading to greater accuracy, better usability, and fewer hallucinations. ### Real-World Applications of Fine-Tuning Fine-tuning can be applied across various domains and needs. Here are a few practical examples of how it makes a difference: \* \*\*Sentiment Analysis for Finance\*\* – Train an LLM to determine if a news headline impacts a company positively or negatively, tailoring its understanding to financial context. \* \*\*Customer Support Chatbots\*\* – Fine-tune on past customer interactions to provide more accurate and personalized responses in a company’s style and terminology. \* \*\*Legal Document Assistance\*\* – Fine-tune on legal texts (contracts, case law, regulations) for tasks like contract analysis, case law research, or compliance support, ensuring the model uses precise legal language. ## The Benefits of Fine-Tuning Fine-tuning offers several notable benefits beyond what a base model or a purely retrieval-based system can provide: #### Fine-Tuning vs. RAG: What’s the Difference? Fine-tuning can do mostly everything RAG can - but not the other way around. During training, fine-tuning embeds external knowledge directly into the model. This allows the model to handle niche queries, summarize documents, and maintain context without relying on an outside retrieval system. That’s not to say RAG lacks advantages as it is excels at accessing up-to-date information from external databases. It is in fact possible to retrieve fresh data with fine-tuning as well, however it is better to combine RAG with fine-tuning for efficiency. #### Task-Specific Mastery Fine-tuning deeply integrates domain knowledge into the model. This makes it highly effective at handling structured, repetitive, or nuanced queries, scenarios where RAG-alone systems often struggle. In other words, a fine-tuned model becomes a specialist in the tasks or content it was trained on. #### Independence from Retrieval A fine-tuned model has no dependency on external data sources at inference time. It remains reliable even if a connected retrieval system fails or is incomplete, because all needed information is already within the model’s own parameters. This self-sufficiency means fewer points of failure in production. #### Faster Responses Fine-tuned models don’t need to call out to an external knowledge base during generation. Skipping the retrieval step means they can produce answers much more quickly. This speed makes fine-tuned models ideal for time-sensitive applications where every second counts. #### Custom Behavior and Tone Fine-tuning allows precise control over how the model communicates. This ensures the model’s responses stay consistent with a brand’s voice, adhere to regulatory requirements, or match specific tone preferences. You get a model that not only knows \*what\* to say, but \*how\* to say it in the desired style. #### Reliable Performance Even in a hybrid setup that uses both fine-tuning and RAG, the fine-tuned model provides a reliable fallback. If the retrieval component fails to find the right information or returns incorrect data, the model’s built-in knowledge can still generate a useful answer. This guarantees more consistent and robust performance for your system. ## Common Misconceptions Despite fine-tuning’s advantages, a few myths persist. Let’s address two of the most common misconceptions about fine-tuning: ### Does Fine-Tuning Add New Knowledge to a Model? \*\*Yes - it absolutely can.\*\* A common myth suggests that fine-tuning doesn’t introduce new knowledge, but in reality it does. If your fine-tuning dataset contains new domain-specific information, the model will learn that content during training and incorporate it into its responses. In effect, fine-tuning \*can and does\* teach the model new facts and patterns from scratch. ### Is RAG Always Better Than Fine-Tuning? \*\*Not necessarily.\*\* Many assume RAG will consistently outperform a fine-tuned model, but that’s not the case when fine-tuning is done properly. In fact, a well-tuned model often matches or even surpasses RAG-based systems on specialized tasks. Claims that “RAG is always better” usually stem from fine-tuning attempts that weren’t optimally configured - for example, using incorrect \[LoRA parameters\](/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide.md) or insufficient training. Unsloth takes care of these complexities by automatically selecting the best parameter configurations for you. All you need is a good-quality dataset, and you'll get a fine-tuned model that performs to its fullest potential. ### Is Fine-Tuning Expensive? \*\*Not at all!\*\* While full fine-tuning or pretraining can be costly, these are not necessary (pretraining is especially not necessary). In most cases, LoRA or QLoRA fine-tuning can be done for minimal cost. In fact, with Unsloth’s \[free notebooks\](https://docs.unsloth.ai/get-started/unsloth-notebooks) for Colab or Kaggle, you can fine-tune models without spending a dime. Better yet, you can even fine-tune locally on your own device. ## FAQ: ### Why You Should Combine RAG & Fine-Tuning Instead of choosing between RAG and fine-tuning, consider using \*\*both\*\* together for the best results. Combining a retrieval system with a fine-tuned model brings out the strengths of each approach. Here’s why: \* \*\*Task-Specific Expertise\*\* – Fine-tuning excels at specialized tasks or formats (making the model an expert in a specific area), while RAG keeps the model up-to-date with the latest external knowledge. \* \*\*Better Adaptability\*\* – A fine-tuned model can still give useful answers even if the retrieval component fails or returns incomplete information. Meanwhile, RAG ensures the system stays current without requiring you to retrain the model for every new piece of data. \* \*\*Efficiency\*\* – Fine-tuning provides a strong foundational knowledge base within the model, and RAG handles dynamic or quickly-changing details without the need for exhaustive re-training from scratch. This balance yields an efficient workflow and reduces overall compute costs. ### LoRA vs. QLoRA: Which One to Use? When it comes to implementing fine-tuning, two popular techniques can dramatically cut down the compute and memory requirements: \*\*LoRA\*\* and \*\*QLoRA\*\*. Here’s a quick comparison of each: \* \*\*LoRA (Low-Rank Adaptation)\*\* – Fine-tunes only a small set of additional “adapter” weight matrices (in 16-bit precision), while leaving most of the original model unchanged. This significantly reduces the number of parameters that need updating during training. \* \*\*QLoRA (Quantized LoRA)\*\* – Combines LoRA with 4-bit quantization of the model weights, enabling efficient fine-tuning of very large models on minimal hardware. By using 4-bit precision where possible, it dramatically lowers memory usage and compute overhead. We recommend starting with \*\*QLoRA\*\*, as it’s one of the most efficient and accessible methods available. Thanks to Unsloth’s \[dynamic 4-bit\](https://unsloth.ai/blog/dynamic-4bit) quants, the accuracy loss compared to standard 16-bit LoRA fine-tuning is now negligible. ### Experimentation is Key There’s no single “best” approach to fine-tuning - only best practices for different scenarios. It’s important to experiment with different methods and configurations to find what works best for your dataset and use case. A great starting point is \*\*QLoRA (4-bit)\*\*, which offers a very cost-effective, resource-friendly way to fine-tune models without heavy computational requirements. {% content-ref url="/pages/y6obKRSk8TwyjIrCjuGE" %} \[Hyperparameters Guide\](/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide.md) {% endcontent-ref %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/faq-+-is-fine-tuning-right-for-me.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/inference-and-deployment/troubleshooting-inference.md). # Troubleshooting Inference ### Running in Unsloth works well, but after exporting & running on other platforms, the results are poor You might sometimes encounter an issue where your model runs and produces good results on Unsloth, but when you use it on another platform like Ollama or vLLM, the results are poor or you might get gibberish, endless/infinite generations \*or\* repeated outputs\*\*.\*\* \* The most common cause of this error is using an \*\*incorrect chat template\*\*\*\*.\*\* It’s essential to use the SAME chat template that was used when training the model in Unsloth and later when you run it in another framework, such as llama.cpp or Ollama. When inferencing from a saved model, it's crucial to apply the correct template. \* You must use the correct \`eos token\`. If not, you might get gibberish on longer generations. \* It might also be because your inference engine adds an unnecessary "start of sequence" token (or the lack of thereof on the contrary) so ensure you check both hypotheses! \* \*\*Use our conversational notebooks to force the chat template - this will fix most issues.\*\* \* Qwen-3 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* Gemma-3 4B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) \* Llama-3.2 3B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* Phi-4 14B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* Mistral v0.3 7B Conversational notebook \[\*\*Open in Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \*\*More notebooks in our\*\* \[\*\*notebooks repo\*\*\](https://github.com/unslothai/notebooks)\*\*.\*\* ### Saving to \`safetensors\`, not \`bin\` format in Colab We save to \`.bin\` in Colab so it's like 4x faster, but set \`safe\_serialization = None\` to force saving to \`.safetensors\`. So \`model.save\_pretrained(..., safe\_serialization = None)\` or \`model.push\_to\_hub(..., safe\_serialization = None)\` ### If saving to GGUF or vLLM 16bit crashes You can try reducing the maximum GPU usage during saving by changing \`maximum\_memory\_usage\`. The default is \`model.save\_pretrained(..., maximum\_memory\_usage = 0.75)\`. Reduce it to say 0.5 to use 50% of GPU peak memory or lower. This can reduce OOM crashes during saving. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/inference-and-deployment/troubleshooting-inference.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/updating.md). # Updating Unsloth ### \*\*Update Unsloth Studio\*\* You can use the same install commands to update #### \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` #### \*\*Windows PowerShell:\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` ### Updating Unsloth Core: \`\`\`bash pip install --upgrade unsloth unsloth\_zoo \`\`\` #### Updating Unsloth Core without dependency updates: pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth_zoo \#### To use an old version of Unsloth: {% code overflow="wrap" %} \`\`\`bash pip install --force-reinstall --no-cache-dir --no-deps unsloth==2025.1.5 \`\`\` {% endcode %} '2025.1.5' is one of the previous old versions of Unsloth. Change it to a specific release listed on our \[Github here\](https://github.com/unslothai/unsloth/releases). --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/updating.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/rl-reward-hacking.md). # RL Reward Hacking The ultimate goal of RL is to maximize some reward (say speed, revenue, some metric). But RL can \*\*cheat.\*\* When the RL algorithm learns a trick or exploits something to increase the reward, without actually doing the task at end, this is called "\*\*Reward Hacking\*\*". It's the reason models learn to modify unit tests to pass coding challenges, and these are critical blockers for real world deployment. Some other good examples are from \[Wikipedia\](https://en.wikipedia.org/wiki/Reward\_hacking). ![](https://i.pinimg.com/originals/55/e0/1b/55e01b94a9c5546b61b59ae300811c83.gif) \*\*Can you counter reward hacking? Yes!\*\* In our \[free gpt-oss RL notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-\\(20B\\)-GRPO.ipynb) we explore how to counter reward hacking in a code generation setting and showcase tangible solutions to common error modes. We saw the model edit the timing function, outsource to other libraries, cache the results, and outright cheat. After countering, the result is our model generates genuinely optimized matrix multiplication kernels, not clever cheats. ## :trophy: Reward Hacking Overview Some common examples of reward hacking during RL include: #### Laziness RL learns to use Numpy, Torch, other libraries, which calls optimized CUDA kernels. We can stop the RL algorithm from calling optimized code by inspecting if the generated code imports other non standard Python libraries. #### Caching & Cheating RL learns to cache the result of the output and RL learns to find the actual output by inspecting Python global variables. We can stop the RL algorithm from using cached data by wiping the cache with a large fake matrix. We also have to benchmark carefully with multiple loops and turns. #### Cheating RL learns to edit the timing function to make it output 0 time as passed. We can stop the RL algorithm from using global or cached variables by restricting it's \`locals\` and \`globals\`. We are also going to use \`exec\` to create the function, so we have to save the output to an empty dict. We also disallow global variable access via \`types.FunctionType(f.\_\_code\_\_, {})\`\\\\ --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/advanced-rl-documentation/rl-reward-hacking.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install/google-colab.md). # Google Colab ![](https://unsloth.ai/files/dIMS0kS50z6kdR2yCJJp) If you have never used a Colab notebook, a quick primer on the notebook itself: 1. \*\*Play Button at each "cell".\*\* Click on this to run that cell's code. You must not skip any cells and you must run every cell in chronological order. If you encounter errors, simply rerun the cell you did not run. Another option is to click CTRL + ENTER if you don't want to click the play button. 2. \*\*Runtime Button in the top toolbar.\*\* You can also use this button and hit "Run all" to run the entire notebook in 1 go. This will skip all the customization steps, but is a good first try. 3. \*\*Connect / Reconnect T4 button.\*\* T4 is the free GPU Google is providing. It's quite powerful! The first installation cell looks like below: Remember to click the PLAY button in the brackets \\\[ \]. We grab our open source Github package, and install some other packages. ![](https://unsloth.ai/files/z9rRDRG3W2lZ6cw69z68) \### Colab Example Code Unsloth example code to fine-tune gpt-oss-20b: \`\`\`python from unsloth import FastLanguageModel, FastModel import torch from trl import SFTTrainer, SFTConfig from datasets import load\_dataset max\_seq\_length = 2048 # Supports RoPE Scaling internally, so choose any! # Get LAION dataset url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified\_chip2.jsonl" dataset = load\_dataset("json", data\_files = {"train" : url}, split = "train") # 4bit pre quantized models we support for 4x faster downloading + no OOMs. fourbit\_models = \[ "unsloth/gpt-oss-20b-unsloth-bnb-4bit", #or choose any model \] # More models at https://huggingface.co/unsloth model, tokenizer = FastModel.from\_pretrained( model\_name = "unsloth/gpt-oss-20b", max\_seq\_length = 2048, # Choose any for long context! load\_in\_4bit = True, # 4-bit quantization. False = 16-bit LoRA. load\_in\_8bit = False, # 8-bit quantization load\_in\_16bit = False, # \[NEW!\] 16-bit LoRA full\_finetuning = False, # Use for full fine-tuning. # token = "hf\_...", # use one if using gated models ) # Do model patching and add fast LoRA weights model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\], lora\_alpha = 16, lora\_dropout = 0, # Supports any, but = 0 is optimized bias = "none", # Supports any, but = "none" is optimized # \[NEW\] "unsloth" uses 30% less VRAM, fits 2x larger batch sizes! use\_gradient\_checkpointing = "unsloth", # True or "unsloth" for very long context random\_state = 3407, max\_seq\_length = max\_seq\_length, use\_rslora = False, # We support rank stabilized LoRA loftq\_config = None, # And LoftQ ) trainer = SFTTrainer( model = model, train\_dataset = dataset, tokenizer = tokenizer, args = SFTConfig( max\_seq\_length = max\_seq\_length, per\_device\_train\_batch\_size = 2, gradient\_accumulation\_steps = 4, warmup\_steps = 10, max\_steps = 60, logging\_steps = 1, output\_dir = "outputs", optim = "adamw\_8bit", seed = 3407, ), ) trainer.train() # Go to https://docs.unsloth.ai for advanced tips like # (1) Saving to GGUF / merging to 16bit for vLLM # (2) Continued training from a saved LoRA adapter # (3) Adding an evaluation loop / OOMs # (4) Customized chat templates \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install/google-colab.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/new/studio/export.md). # Export models with Unsloth Studio Use \[Unsloth Studio\](/docs/new/studio.md) to export, save, or convert models to GGUF, Safetensors, or LoRA for deployment, sharing, or local inference in Unsloth, llama.cpp, Ollama, vLLM, and more. Export a trained checkpoint or convert any existing model. ![](https://unsloth.ai/files/DuavmU7CXLCHCQ9cQjsX) {% stepper %} {% step %} ### Select Training Run Start by selecting the training run you want to export from. Each run represents a complete training session and may contain multiple checkpoints. After choosing a run, select the checkpoint to export. A checkpoint is a saved version of the model created during training. ![](https://unsloth.ai/files/fnJc6KlndZIS9u3R5oaJ) {% endstep %} {% step %} ### Select Checkpoint Later checkpoints typically represent the final trained model, but you can export any checkpoint depending on your needs. ![](https://unsloth.ai/files/Zq26gwGpJyZzABYfrQIZ) {% endstep %} {% step %} ### Export Methods Depending on your workflow, you can export a merged model, LoRA adapter weights, or a GGUF model for local inference. ![](https://unsloth.ai/files/evU02j9aUoFc85GVwmfw) Each export method produces a different version of the model depending on how you plan to run or share it. The table below explains what each option exports. | Export Type | Description | | ---------------- | ------------------------------------------------------------------------------------------------- | | Merged Model | \*\*16-bit model\*\* with the LoRA adapter merged into the base weights. | | LoRA Only | Exports \*\*only the adapter weights\*\*. Requires the original base model. | | GGUF / llama.cpp | Converts the model to \*\*GGUF format\*\* for Unsloth / llama.cpp \*\*/\*\* Ollama / LM Studio inference. | | {% endstep %} | | {% step %} ### Export / Save Locally When exporting a model, you can choose where the resulting files should be saved. Models can be downloaded directly to your machine or pushed to the Hugging Face Hub for hosting and sharing. Save the exported model files directly to your machine. This option is useful for running the model locally, distributing files manually, or integrating with local inference tools. ![](https://unsloth.ai/files/QdBIlnhrCV8E2yjg5p5i) {% endstep %} {% step %} ### Push to Hub Upload the exported model to the Hugging Face Hub. This allows you to host, share, and deploy the model from a central repository. You will need a Hugging Face write token to publish the model. ![](https://unsloth.ai/files/WEVYsNx4rYtLqqfGi0BV) {% hint style="success" %} If you are already authenticated with the Hugging Face CLI, the write token can be left empty. {% endhint %} {% endstep %} {% endstepper %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/new/studio/export.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/openrouter.md). # How to Connect OpenRouter to Unsloth: API Key & Model Setup This guide explains how to connect \*\*OpenRouter to\*\* \[\*\*Unsloth\*\*\](https://github.com/unslothai/unsloth) so you can access hosted AI models from providers like \*\*OpenAI, Anthropic,\*\* and \*\*Google\*\* through an open-source local UI chat interface. You’ll learn how to create an OpenRouter API key, add OpenRouter as a provider in Unsloth, load or manually enter model IDs, and enable external models for chat. Once a single API key is connected, OpenRouter models in Unsloth can provide advanced features such as thinking, web search, tool calling, code execution, and customizable generation settings directly from the chat page. ### Setup {% stepper %} {% step %} #### Create an OpenRouter API key Sign in to your OpenRouter account. Create an API key from the \[OpenRouter dashboard\](https://openrouter.ai/settings/keys): ![](https://unsloth.ai/files/eq9ChRNda6w66bG9BMAH) Copy the key. You will paste it into Unsloth in the next step. When creating the key, you can optionally set a credit limit or expiration date. {% endstep %} {% step %} #### Connect OpenRouter to Unsloth Open \*\*Settings → Connections\*\*, then click \*\*Add Connected\*\*. Select \*\*OpenRouter\*\*, then enter your connection details. Enter your OpenRouter details: \* \*\*API key:\*\* paste your OpenRouter API key ![](https://unsloth.ai/files/LK2es112MrFf1wIg4fo4) \* \*\*Model IDs:\*\* click \*\*Load Models\*\*, or enter model IDs manually ![](https://unsloth.ai/files/3HzfTW7WXVvclpFn5AHy) Finally, click \*\*Add Connection\*\*. {% endstep %} {% step %} #### Ready to Chat After saving the connection, select an OpenRouter model under \*\*Connected\*\* in the model dropdown. ![](https://unsloth.ai/files/As8Bp6CzxvzU7YVRI4bx) OpenRouter models can expose different controls depending on the upstream model, including web search, thinking, tool-calling, and generation settings. {% endstep %} {% endstepper %} ### Model Selection OpenRouter provides access to many models from different providers. If \*\*Load Models\*\* does not return the models you want to select, enter the model IDs you want enabled. ![](https://unsloth.ai/files/mUoNuWesdwHDpWYzZ6RG) Example model IDs: \`\`\` openai/gpt-5.5 anthropic/claude-sonnet-4.6 google/gemini-3-pro \`\`\` ### Troubleshooting If OpenRouter fails to connect, check that the API key is valid and belongs to the correct OpenRouter account. If a model does not appear after clicking \*\*Load Models\*\*, it may not be available for your account or region. You can enter the model ID manually or choose another model. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/openrouter.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/integrations/connections/anthropic-claude.md). # Connect Anthropic to Unsloth: Run Claude Models in Local Chat Connect the Anthropic API to \[Unsloth\](https://github.com/unslothai/unsloth) to chat with Claude models including Claude Opus 4.7 directly alongside your local models in an open-source UI chat interface. This guide shows you how to create an Anthropic API key, add Anthropic as a provider in Unsloth, load Claude LLMs, and start chatting. Supported Claude models in Unsloth can also access advanced features such as thinking, \[web search\](#web-search-and-thinking), Anthropic \[code execution\](#code-execution), and \[prompt caching\](#prompt-caching) to improve cost effiencey. ### Setup {% stepper %} {% step %} #### Create an Anthropic API key Create an API key from the \[Anthropic Console\](https://console.anthropic.com/settings/keys). Copy the key. You will paste it into Unsloth in the next step. {% endstep %} {% step %} #### Connect Anthropic to Unsloth ![](https://unsloth.ai/files/4lIKuZpaZFJ16VCXeyAH) Next, connect Anthropic to Unsloth. 1. Open \*\*Settings\*\* → \*\*Connections\*\*, then click \*\*Add Connection.\*\* 2. Select the provider you want to add, then paste the API key you copied earlier. 3. Click \*\*Reload Models\*\* to refresh the list with models available to your account. 4. Choose the models you want to enable, then hit save. {% endstep %} {% step %} #### Ready to Chat After saving the connection, select a Claude model under \*\*Connected\*\* in the model dropdown. Supported Claude models can expose extra controls including image generation, thinking, web search and code execution. ![](https://unsloth.ai/files/j7XVwXmv37LeP0VWlu7g) {% endstep %} {% endstepper %} ### Code Execution When enabled, supported Claude models can run code in Anthropic’s provider sandbox to solve problems, analyze data, and work with files. Claude uses Anthropic’s Code execution tool. Code execution appears in the response timeline as tool activity, alongside other tool calls. ![](https://unsloth.ai/files/3iL48hdRoVyoL73z8DP8) \### Prompt Caching {% columns %} {% column width="66.66666666666666%" %} Prompt caching reduces latency and cost when requests reuse the same long prefix. It is supported for compatible providers and servers, including Anthropic models. Use the \*\*Prompt caching\*\* setting in the side panel to control caching behavior for supported connections. {% endcolumn %} {% column width="33.33333333333334%" %} ![](https://unsloth.ai/files/QSJQ82qfp3z5H35zIou6) {% endcolumn %} {% endcolumns %} ![](https://unsloth.ai/files/rd7uqDkUz6YnddW01aRl) \### Web Search & Thinking Supported Claude models can use provider-side web search. The \*\*Think\*\* control appears when the selected model supports thinking. Depending on the model, this may expose different thinking levels or availability. ![](https://unsloth.ai/files/46johOqXvwqWOhGnbso4) The \*\*Think\*\* control adapts to the selected model: some models use an on/off toggle, while reasoning-effort models use model specific thinking levels. ### Troubleshooting If Anthropic API fails to connect, check that the API key is valid and belongs to the correct Anthropic account. If a model does not appear after clicking \*\*Load Models\*\*, it may not be available for your account. You can enter the model ID manually or choose another model. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/integrations/connections/anthropic-claude.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/get-started/install.md). # Unsloth Installation Unsloth can be used in two ways: through \[Unsloth Studio\](/docs/new/studio/install.md), the web UI, or through Unsloth Core, the original code-based version. See our \[system requirements\](/docs/get-started/fine-tuning-for-beginners/unsloth-requirements.md) Unsloth Studio works on MacOS, Linux, Windows, NVIDIA, and more. Use the same install commands to update. \*\*MacOS, Linux, WSL:\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell:\*\* \`\`\`powershell irm https://unsloth.ai/install.ps1 | iex \`\`\` \*\*Launch Unsloth Studio:\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` | | Cover image | | | --- | --- | --- | | [/pages/nQJglux1e9VKfVL4F43M](https://unsloth.ai/pages/nQJglux1e9VKfVL4F43M) | | | | [/pages/LhZlVJv6yLKmWbNy4UmP](https://unsloth.ai/pages/LhZlVJv6yLKmWbNy4UmP) | | [/pages/LhZlVJv6yLKmWbNy4UmP](https://unsloth.ai/pages/LhZlVJv6yLKmWbNy4UmP) | | [/pages/Sv0QKHkAGvwTK47OKX2A](https://unsloth.ai/pages/Sv0QKHkAGvwTK47OKX2A) | | | | [/pages/GUxDmG8LaAiQinraCXCr](https://unsloth.ai/pages/GUxDmG8LaAiQinraCXCr) | | | | [/pages/SQoCZEpSeGsxtIypEgup](https://unsloth.ai/pages/SQoCZEpSeGsxtIypEgup) | | [/pages/SQoCZEpSeGsxtIypEgup](https://unsloth.ai/pages/SQoCZEpSeGsxtIypEgup) | | [/pages/dZaYUyA34oYX3LotyGAB](https://unsloth.ai/pages/dZaYUyA34oYX3LotyGAB) | | | | [/pages/FhVmcV9yU5zmKQvNYNb8](https://unsloth.ai/pages/FhVmcV9yU5zmKQvNYNb8) | | | | [/pages/nuQrnhqDQWCXcjxOPAKP](https://unsloth.ai/pages/nuQrnhqDQWCXcjxOPAKP) | | [/pages/nuQrnhqDQWCXcjxOPAKP](https://unsloth.ai/pages/nuQrnhqDQWCXcjxOPAKP) | | [/pages/oyA2PFAEd9nf2dQXnWxq](https://unsloth.ai/pages/oyA2PFAEd9nf2dQXnWxq) | | | | [/pages/tyk0jWHCp0YFpzPIgIUk](https://unsloth.ai/pages/tyk0jWHCp0YFpzPIgIUk) | | [/pages/tyk0jWHCp0YFpzPIgIUk](https://unsloth.ai/pages/tyk0jWHCp0YFpzPIgIUk) | \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/get-started/install.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth.md). # Fine-tuning LLMs with NVIDIA DGX Spark and Unsloth Unsloth enables local fine-tuning of LLMs with up to \*\*200B parameters\*\* on the NVIDIA DGX™ Spark. With 128 GB of unified memory, you can train massive models such as \*\*gpt-oss-120b\*\*, and run or deploy inference directly on DGX Spark. As shown at \[OpenAI DevDay\](https://x.com/UnslothAI/status/1976284209842118714), gpt-oss-20b was trained with RL and Unsloth on DGX Spark to auto-win 2048. You can train using Unsloth in a Docker container or virtual environment on DGX Spark. ![](https://unsloth.ai/files/v6KqlW8eqrrmnU0yBCu1) ![](https://unsloth.ai/files/WR7Bq46fl2QipCM3fJPV) In this tutorial, we’ll train gpt-oss-20b with RL using Unsloth notebooks after installing Unsloth on your DGX Spark. gpt-oss-120b will use around \*\*68GB\*\* of unified memory. After 1,000 steps and 4 hours of RL training, the gpt-oss model greatly outperforms the original on 2048, and longer training would further improve results. ![](https://unsloth.ai/files/nkAjEpiLMJUnQZD9qyLi) You can watch Unsloth featured on OpenAI DevDay 2025 [here](https://youtu.be/1HL2YHRj270?si=8SR6EChF34B1g-5r&t=1080) . ![](https://unsloth.ai/files/GH7Z2Fi2ad5ptgpRU5t9) gpt-oss trained with RL consistently outperforms on 2048. \### ⚡ Step-by-Step Tutorial {% stepper %} {% step %} \*\*Start with Unsloth Docker image for DGX Spark\*\* First, build the Docker image using the DGX Spark Dockerfile which can be \[found here\](https://raw.githubusercontent.com/unslothai/notebooks/main/Dockerfile\_DGX\_Spark). You can also run the below in a Terminal in the DGX Spark: \`\`\`bash sudo apt update && sudo apt install -y wget wget -O Dockerfile "https://raw.githubusercontent.com/unslothai/notebooks/main/Dockerfile\_DGX\_Spark" \`\`\` Then, build the training Docker image using saved Dockerfile: \`\`\`bash docker build -f Dockerfile -t unsloth-dgx-spark . \`\`\` ![](https://unsloth.ai/files/M0g66UIJFrA5GmN5vRkx) You can also click to see the full DGX Spark Dockerfile \`\`\`python FROM nvcr.io/nvidia/pytorch:25.09-py3 # Set CUDA environment variables ENV CUDA\_HOME=/usr/local/cuda-13.0/ ENV CUDA\_PATH=$CUDA\_HOME ENV PATH=$CUDA\_HOME/bin:$PATH ENV LD\_LIBRARY\_PATH=$CUDA\_HOME/lib64:$LD\_LIBRARY\_PATH ENV C\_INCLUDE\_PATH=$CUDA\_HOME/include:$C\_INCLUDE\_PATH ENV CPLUS\_INCLUDE\_PATH=$CUDA\_HOME/include:$CPLUS\_INCLUDE\_PATH # Install triton from source for latest blackwell support RUN git clone https://github.com/triton-lang/triton.git && \\ cd triton && \\ git checkout c5d671f91d90f40900027382f98b17a3e04045f6 && \\ pip install -r python/requirements.txt && \\ pip install . && \\ cd .. # Install xformers from source for blackwell support RUN git clone --depth=1 https://github.com/facebookresearch/xformers --recursive && \\ cd xformers && \\ export TORCH\_CUDA\_ARCH\_LIST="12.1" && \\ python setup.py install && \\ cd .. # Install unsloth and other dependencies RUN pip install unsloth unsloth\_zoo bitsandbytes==0.48.0 transformers==4.56.2 trl==0.22.2 # Launch the shell CMD \["/bin/bash"\] \`\`\` {% endstep %} {% step %} \*\*Launch container\*\* Launch the training container with GPU access and volume mounts: \`\`\`bash docker run -it \\ --gpus=all \\ --net=host \\ --ipc=host \\ --ulimit memlock=-1 \\ --ulimit stack=67108864 \\ -v $(pwd):$(pwd) \\ -v $HOME/.cache/huggingface:/root/.cache/huggingface \\ -w $(pwd) \\ unsloth-dgx-spark \`\`\` ![](https://unsloth.ai/files/8WyYpnaFGgaWlryPgFjB) ![](https://unsloth.ai/files/kFY0n1pYzi5Te61UpLgF) {% endstep %} {% step %} \*\*Start Jupyter and Run Notebooks\*\* Inside the container, start Jupyter and run the required notebook. You can use the Reinforcement Learning gpt-oss 20b to win 2048 \[notebook here\](https://github.com/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_Reinforcement\_Learning\_2048\_Game\_DGX\_Spark.ipynb). In fact all \[Unsloth notebooks\](https://docs.unsloth.ai/get-started/unsloth-notebooks) work in DGX Spark including the \*\*120b\*\* notebook! Just remove the installation cells. ![](https://unsloth.ai/files/WR7Bq46fl2QipCM3fJPV) The below commands can be used to run the RL notebook as well. After Jupyter Notebook is launched, open up the “\`gpt\_oss\_20B\_RL\_2048\_Game.ipynb\`” \`\`\`bash NOTEBOOK\_URL="https://raw.githubusercontent.com/unslothai/notebooks/refs/heads/main/nb/gpt\_oss\_(20B)\_Reinforcement\_Learning\_2048\_Game\_DGX\_Spark.ipynb" wget -O "gpt\_oss\_20B\_RL\_2048\_Game.ipynb" "$NOTEBOOK\_URL" jupyter notebook --ip=0.0.0.0 --port=8888 --no-browser --allow-root \`\`\` ![](https://unsloth.ai/files/Vwd0Jo2pN8NQcTYUpUPb) Don't forget Unsloth also allows you to \[save and run\](/docs/basics/inference-and-deployment.md) your models after fine-tuning so you can locally deploy them directly on your DGX Spark after. {% endstep %} {% endstepper %} Many thanks to \[Lakshmi Ramesh\](https://www.linkedin.com/in/rlakshmi24/) and \[Barath Anandan\](https://www.linkedin.com/in/barathsa/) from NVIDIA for helping Unsloth’s DGX Spark launch and building the Docker image. ### Unified Memory Usage gpt-oss-120b QLoRA 4-bit fine-tuning will use around \*\*68GB\*\* of unified memory. How your unified memory usage should look \*\*before\*\* (left) and \*\*after\*\* (right) training: ![](https://unsloth.ai/files/V1zEMZXvKKj4A5mkWKgK) ![](https://unsloth.ai/files/Kucl1Tmep9YD2nOY5aLL) And that's it! Have fun training and running LLMs completely locally on your NVIDIA DGX Spark! ### Video Tutorials Thanks to Tim from \[AnythingLLM\](https://github.com/Mintplex-Labs/anything-llm) for providing a great fine-tuning tutorial with Unsloth on DGX Spark: {% embed url="" %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/fine-tuning-llms-with-nvidia-dgx-spark-and-unsloth.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/500k-context-length-fine-tuning.md). # 500K Context Length Fine-tuning We’re introducing new algorithms in Unsloth that push the limits of long-context training for \*\*any LLM and VLM\*\*. Training LLMs like gpt-oss-20b can now reach \*\*500K+ context lengths\*\* on a single 80GB H100 GPU, compared to 80K previously with no accuracy degradation. You can reach >\*\*750K context windows\*\* on a B200 192GB GPU. > \*\*Try 500K-context gpt-oss-20b fine-tuning on our\*\* \[\*\*80GB A100 Colab notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_500K\_Context\_Fine\_tuning.ipynb)\*\*.\*\* We’ve significantly improved how Unsloth handles memory usage patterns, speed, and context lengths: \* \*\*60% lower VRAM use\*\* with \*\*3.2x longer context\*\* via Unsloth’s new \[fused and chunked cross-entropy\](#unsloth-loss-refactoring-chunk-and-fuse) loss, with no degradation in speed or accuracy \* Enhanced activation offloading in Unsloth’s \[\*\*Gradient Checkpointing\*\*\](#unsloth-gradient-checkpointing-enhanced) \* Collabing with Stas Bekman from Snowflake on \[Tiled MLP\](#tiled-mlp-unlocking-500k), enabling 2× more contexts Unsloth’s algorithms allows gpt-oss-20b QLoRA (4bit) with 290K context possible on a H100 with no accuracy loss, and 500K+ with Tiled MLP enabled, altogether delivering >\*\*6.4x longer context lengths.\*\* ![](https://unsloth.ai/files/0aIvZ3MYsPhPDGuOPyoC) \### 📐 Unsloth Loss Refactoring: Chunk & Fuse Our new fused loss implementation adds \*\*dynamic sequence chunking\*\*: instead of computing language model head logits and cross-entropies over the entire sequence at once, we process manageable slices along the flattened sequence dimension. This cuts peak memory from GBs to a smaller chunk sizes. Each chunk still runs a fully fused forward + backward pass via \`torch.func.grad\_and\_value\` , and retains mixed precision accuracy by upcasting to float32 if necessary. \*\*These changes do not degrade training speed or accuracy.\*\* ![](https://unsloth.ai/files/m28CeGE9I9POS7nRblCt) The key innovation is that the \*\*chunk size is chosen automatically at runtime\*\* based on available VRAM. \* If you have more free VRAM, larger chunks are used for faster runs \* If you have less VRAM, it increases the number of chunks to avoid memory blowouts. This \*\*removes manual tuning\*\* and keeps our algorithm robust across old and new GPUs, workloads and different sequence lengths. {% hint style="success" %} Due to automatic tuning, \*\*smaller contexts will use more VRAM\*\* (fewer chunks) to \*\*avoid unnecessary overhead\*\*. For the plots above, we adjust the number of loss chunks to reflect realistic VRAM tiers. With 80GB VRAM, this yields >3.2× longer contexts. {% endhint %} ### 🏁 Unsloth Gradient Checkpointing Enhancements Our \[Unsloth Gradient Checkpointing\](https://unsloth.ai/blog/long-context) algorithm, \*\*introduced in April 2024\*\*, quickly became popular and the standard across the industry, having been integrated into most training packages nowadays. It offloads activations to CPU RAM which allowed 10x longer context lengths. Our new enhancements uses CUDA Streams and other tricks to add at most \*\*0.1%\*\* training overhead with no impact on accuracy. Previously it added 1 to 3% training overhead. {% code expandable="true" %} \`\`\`python # Original Unsloth version released April 2024 - LGPLv3 Licensed class Unsloth\_Offloaded\_Gradient\_Checkpointer(torch.autograd.Function): @staticmethod @torch\_amp\_custom\_fwd def forward(ctx, forward\_function, hidden\_states, \*args): ctx.device = hidden\_states.device saved\_hidden\_states = hidden\_states.to("cpu", non\_blocking = True) with torch.no\_grad(): output = forward\_function(hidden\_states, \*args) ctx.save\_for\_backward(saved\_hidden\_states) ctx.forward\_function, ctx.args = forward\_function, args return output @staticmethod @torch\_amp\_custom\_bwd def backward(ctx, dY): (hidden\_states,) = ctx.saved\_tensors hidden\_states = hidden\_states.to(ctx.device, non\_blocking = True).detach() hidden\_states.requires\_grad\_(True) with torch.enable\_grad(): (output,) = ctx.forward\_function(hidden\_states, \*ctx.args) torch.autograd.backward(output, dY) return (None, hidden\_states.grad,) + (None,)\*len(ctx.args) \`\`\` {% endcode %} By offloading activations as soon as they are produced, we minimize peak activation footprint and free GPU memory exactly when it’s needed. This sharply reduces memory pressure in long-context or large-batch training, where a single decoder layer’s activations can exceed 2 GB. > \*\*Thus, Unsloth’s new algorithms & Gradient Checkpointing contributes to most improvements (3.2x), enabling 290k-context-length QLoRA GPT-OSS fine-tuning on a single H100.\*\* ### 🔓 Tiled MLP: Unlocking 500K+ With help from \[Stas Bekman\](https://x.com/StasBekman) (Snowflake), we integrated Tiled MLP from Snowflake’s Arctic Long Sequence Training \[paper\](https://arxiv.org/abs/2506.13996) and blog post. TiledMLP reduces activation memory and enables much longer sequence lengths by tiling hidden states along the sequence dimension before heavy MLP projections. \*\*We also introduce a few quality-of-life improvements:\*\* We preserve RNG state across tiled forward recomputations so dropout and other stochastic ops are consistent between forward and backward replays. This keeps nested checkpointed computations stable and numerically identical. {% hint style="success" %} Our implementation auto patches any module named or typed as \`mlp\`, so \*\*nearly all models with MLP modules are supported out of the box for Tiled MLP.\*\* {% endhint %} \*\*Tradeoffs to keep in mind\*\* TiledMLP saves VRAM at the cost of extra forward passes. Because it lives inside a checkpointed transformer block and is itself written in a checkpoint style, it effectively becomes a nested checkpoint: one \*\*MLP now performs \\~3 forward passes and 1 backward pass per step\*\*. In return, we can drop almost all intermediate MLP activations from VRAM while still supporting extremely long sequences. ![](https://unsloth.ai/files/RARK9HOzrj33vHjZJpcv) The plots compare active memory timelines for a single decoder layer’s forward and backward during a long-context training step, without Tiled MLP (left) and with it (right). Without Tiled MLP, peak VRAM occurs during the MLP backward; with Tiled MLP, it shifts to the fused loss calculation. We see \\~40% lower VRAM usage, and because the fused loss auto chunks dynamically based on available VRAM, the peak with Tiled MLP would be even smaller on smaller GPUs. ![](https://unsloth.ai/files/JsHzsH7Jc199iJFKS3Gh) To show cross-entropy loss is not the new bottleneck, we fix its chunk size instead of choosing it dynamically and then double the number of chunks. This significantly reduces the loss-related memory spikes. The max memory now occurs during backward in both cases, and overall timing is similar, though Tiled MLP adds a small overhead: one large GEMM becomes many sequential matmuls, plus the extra forward pass mentioned above. Overall, the trade-off is worth it: without Tiled MLP, long-context training can require roughly 2× the memory usage, while with \*\*Tiled MLP a single GPU pays only about a 1.3× increase in step time for the same context length.\*\* \*\*Enabling Tiled MLP in Unsloth:\*\* \`\`\`py model, tokenizer = FastLanguageModel.from\_pretrained( ..., unsloth\_tiled\_mlp = True, ) \`\`\` Just set \`unsloth\_tiled\_mlp = True\` in \`from\_pretrained\` and Tiled MLP is enabled. We follow the same logic as the Arctic paper and choose \`num\_shards = ceil(seq\_len/hidden\_size)\`. Each tile will operate on sequence lengths which are the same size of the hidden dimension of the model to balance throughput and memory savings. We also discussed how Tiled MLP effectively does 3 forward passes and 1 backward, compared to normal gradient checkpointing which does 2 forward passes and 1 backward with Stas Bekman and \[DeepSpeed\](https://github.com/deepspeedai/DeepSpeed/pull/7664) provided a doc update for Tiled MLP within DeepSpeed. {% hint style="success" %} Next time fine-tuning runs out of memory, try turning on \`unsloth\_tiled\_mlp = True\`. This should save some VRAM as long as the context length is longer than the LLM's hidden dimension. {% endhint %} \*\*\* \*\*With our latest update, it is possible to now reach 1M context length with a smaller model on a single GPU!\*\* \*\*Try 500K-context gpt-oss-20b fine-tuning on our\*\* \[\*\*80GB A100 Colab notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt\_oss\_\\(20B\\)\_500K\_Context\_Fine\_tuning.ipynb)\*\*.\*\* If you've made it this far, we're releasing a new blog on our latest improvements in training speed this week so stay tuned by joining our \[Reddit r/unsloth\](https://www.reddit.com/r/unsloth/) or our Docs. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/500k-context-length-fine-tuning.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp.md). # Multi-GPU Fine-tuning with Distributed Data Parallel (DDP) Let’s assume we have multiple GPUs, and we want to fine-tune a model using all of them! To do so, the most straightforward strategy is to use Distributed Data Parallel (DDP), which creates one copy of the model on each GPU device, feeding each copy distinct samples from the dataset during training and aggregating their contributions to weight updates per optimizer step. Why would we want to do this? Well, as we add more GPUs into the training process, we scale the number of samples our models train on per step, making each gradient update more stable and increasing our training throughput dramatically with each added GPU. Here’s a step-by-step guide on how to do this using Unsloth’s command-line interface (CLI)! \*\*Note:\*\* Unsloth DDP will work with any of your training scripts, not just via our CLI! More details below. #### Install Unsloth from source We’ll clone Unsloth from GitHub and install it. Please consider using a \[virtual environment\](https://docs.python.org/3/tutorial/venv.html); we like to use \`uv venv –python 3.12 && source .venv/bin/activate\`, but any virtual environment creation tooling will do. \`\`\`bash git clone https://github.com/unslothai/unsloth.git cd unsloth pip install . \`\`\` #### Choose target model and dataset for finetuning In this demo, we will fine-tune \[Qwen/Qwen3-8B\](https://huggingface.co/Qwen/Qwen3-8B) on the \[yahma/alpaca-cleaned\](https://huggingface.co/datasets/yahma/alpaca-cleaned) chat dataset. This is a Supervised Fine-Tuning (SFT) workload that is commonly used when attempting to adapt a base model to a desired conversational style, or improve the model’s performance on a downstream task. ### Use the Unsloth CLI! First, let’s take a look at the help message built-in to the CLI (we’ve abbreviated here with “...” in various places for brevity): {% code expandable="true" %} \`\`\`bash $ python unsloth-cli.py --help usage: unsloth-cli.py \[-h\] \[--model\_name MODEL\_NAME\] \[--max\_seq\_length MAX\_SEQ\_LENGTH\] \[--dtype DTYPE\] \[--load\_in\_4bit\] \[--dataset DATASET\] \[--r R\] \[--lora\_alpha LORA\_ALPHA\] \[--lora\_dropout LORA\_DROPOUT\] \[--bias BIAS\] \[--use\_gradient\_checkpointing USE\_GRADIENT\_CHECKPOINTING\] … 🦥 Fine-tune your llm faster using unsloth! options: -h, --help show this help message and exit 🤖 Model Options: --model\_name MODEL\_NAME Model name to load --max\_seq\_length MAX\_SEQ\_LENGTH Maximum sequence length, default is 2048. We auto support RoPE Scaling internally! … 🧠 LoRA Options: These options are used to configure the LoRA model. --r R Rank for Lora model, default is 16. (common values: 8, 16, 32, 64, 128) --lora\_alpha LORA\_ALPHA LoRA alpha parameter, default is 16. (common values: 8, 16, 32, 64, 128) … 🎓 Training Options: --per\_device\_train\_batch\_size PER\_DEVICE\_TRAIN\_BATCH\_SIZE Batch size per device during training, default is 2. --per\_device\_eval\_batch\_size PER\_DEVICE\_EVAL\_BATCH\_SIZE Batch size per device during evaluation, default is 4. --gradient\_accumulation\_steps GRADIENT\_ACCUMULATION\_STEPS Number of gradient accumulation steps, default is 4. … \`\`\` {% endcode %} This should give you a sense of what options are available for you to pass into the CLI for training your model! For multi-GPU training (DDP in this case), we will use the \[torchrun\](https://docs.pytorch.org/docs/stable/elastic/run.html) launcher, which allows you to spin up multiple distributed training processes in single-node or multi-node settings. In our case, we will focus on the single-node (i.e., one machine) case with two H100 GPUs. Let’s also check our GPUs’ status by using the \`nvidia-smi\` command-line tool: {% code expandable="true" %} \`\`\`bash $ nvidia-smi Mon Nov 24 12:53:00 2025 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA H100 80GB HBM3 On | 00000000:04:00.0 Off | 0 | | N/A 32C P0 69W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ | 1 NVIDIA H100 80GB HBM3 On | 00000000:05:00.0 Off | 0 | | N/A 30C P0 68W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | No running processes found | +-----------------------------------------------------------------------------------------+ \`\`\` {% endcode %} Great! We have two H100 GPUs, as expected. Both are sitting at 0MiB memory usage as we’re currently not training anything, or have any model loaded into memory. To start your training run, issue a command like the following: {% code expandable="true" %} \`\`\`bash # required: # --model\_name # --dataset # optional; experiment with these: # --learning\_rate, --max\_seq\_length, --per\_device\_train\_batch\_size, --gradient\_accumulation\_steps, --max\_steps # to save the model at the end of training: # --save\_model torchrun --nproc\_per\_node=2 unsloth-cli.py \\ --model\_name=Qwen/Qwen3-8B \\ --dataset=yahma/alpaca-cleaned \\ --learning\_rate=2e-5 \\ --max\_seq\_length=2048 \\ --per\_device\_train\_batch\_size=1 \\ --gradient\_accumulation\_steps=4 \\ --max\_steps=1000 \\ --save\_model \`\`\` {% endcode %} If you have more GPUs, you may set \`--nproc\_per\_node\` accordingly to utilize them. \*\*Note:\*\* You can use the \`torchrun\` launcher with any of your Unsloth training scripts, including the \[scripts\](https://github.com/unslothai/notebooks/tree/main/python\_scripts) converted from our free Colab notebooks, and DDP will be auto-enabled when training with >1 GPU! Taking a look again at \`nvidia-smi\` while training is in-flight: {% code expandable="true" %} \`\`\`bash $ nvidia-smi Mon Nov 24 12:58:42 2025 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA H100 80GB HBM3 On | 00000000:04:00.0 Off | 0 | | N/A 38C P0 193W / 700W | 18903MiB / 81559MiB | 25% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ | 1 NVIDIA H100 80GB HBM3 On | 00000000:05:00.0 Off | 0 | | N/A 37C P0 199W / 700W | 18905MiB / 81559MiB | 28% Default | | | | Disabled | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | 0 N/A N/A 4935 C ...und/unsloth/.venv/bin/python3 18256MiB | | 0 N/A N/A 4936 C ...und/unsloth/.venv/bin/python3 630MiB | | 1 N/A N/A 4935 C ...und/unsloth/.venv/bin/python3 630MiB | | 1 N/A N/A 4936 C ...und/unsloth/.venv/bin/python3 18258MiB | +-----------------------------------------------------------------------------------------+ \`\`\` {% endcode %} We can see that both GPUs are now using \\~19GB of VRAM per H100 GPU! Inspecting the training logs, we see that we’re able to train at a rate of \\~1.1 iterations/s. This training speed is \\~constant even as we add more GPUs, so our training throughput increases \\~linearly with the number of GPUs! ### Training metrics We ran a few short rank-16 LoRA fine-tunes on \[unsloth/Llama-3.2-1B-Instruct\](https://huggingface.co/unsloth/Llama-3.2-1B-Instruct) on the \[yahma/alpaca-cleaned\](https://huggingface.co/datasets/yahma/alpaca-cleaned) dataset to demonstrate the improved training throughput when using DDP training with multiple GPUs. ![](https://unsloth.ai/files/6bDuiiln6JQlxvheV3fb) The above figure compares training loss between two Llama-3.2-1B-Instruct LoRA fine-tunes over 500 training steps, with single GPU training (pink) vs. multi-GPU DDP training (blue). Notice that the loss curves match in scale and trend, but otherwise are a \*bit\* different, since \*the multi-GPU training processes twice as much training data per step\*. This results in a slightly different training curve with less variability on a step-by-step basis. ![](https://unsloth.ai/files/IulTA79vzfe9YkTrBKTi) The above figure plots training progress for the same two fine-tunes. Notice that the multi-GPU DDP training progresses through an epoch of the training data in half as many steps as single GPU training. This is because each GPU can process a distinct batch (of size \`per\_device\_train\_batch\_size\`) per step. However, the per-step timing for DDP training is slightly slower due to distributed communication for the model weight updates. As you increase the number of GPUs, the training throughput will continue to increase \\~linearly (but with a small, but increasing penalty for the distributed comms). These same loss and training epoch progress behaviors hold for QLoRA fine-tunes, in which we loaded the base models in 4-bit precision in order to save additional GPU memory. This is particularly useful for training large models on limited amounts of GPU VRAM: ![](https://unsloth.ai/files/Ocw3LIxgRt1cbOgWmQV1) Training loss comparison between two Llama-3.2-1B-Instruct QLoRA fine-tunes over 500 training steps, with single GPU training (orange) vs. multi-GPU DDP training (purple). ![](https://unsloth.ai/files/IjkwFDvbBqDYM5VUAokI) Training progress comparison for the same two fine-tunes. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/multi-gpu-training-with-unsloth/ddp.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/basics/embedding-finetuning.md). # Fine-tuning Embedding Models with Unsloth Guide Fine-tuning embedding models can largely improve retrieval and RAG performance on specific tasks. It aligns the model's vectors with your domain and the kind of 'similarity' that matters for your use case, which improves search, RAG, clustering, and recommendations on your data. Example: The headlines “Google launches Pixel 10” and “Qwen releases Qwen3” might be embedded as similar if you’re just labeling both as 'Tech,' but not similar if you’re doing semantic search because they’re about different things. Fine-tuning helps the model make the 'right' kind of similarity for your use case, reducing errors and improving results. \[\*\*Unsloth\*\*\](https://github.com/unslothai/unsloth) now supports training embedding, \*\*classifier\*\*, \*\*BERT\*\*, \*\*reranker\*\* models \[\*\*\\~1.8-3.3x faster\*\*\](#unsloth-benchmarks) with 20% less memory and 2x longer context than other Flash Attention 2 implementations - no accuracy degradation. EmbeddingGemma-300M works on just \*\*3GB VRAM\*\*. You can use your trained \*\*model anywhere\*\*: transformers, LangChain, Ollama, vLLM, llama.cpp etc. Unsloth uses \[SentenceTransformers\](https://github.com/huggingface/sentence-transformers) to support compatible models like Qwen3-Embedding, BERT and more. \*\*Even if there's no notebook or upload, it’s still supported.\*\* \*\*We created free fine-tuning notebooks, with 3 main use-cases:\*\* | \[EmbeddingGemma (300M)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/EmbeddingGemma\_\\(300M\\).ipynb) | \[Qwen3-Embedding 4B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_Embedding\_\\(4B\\).ipynb) • \[0.6B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_Embedding\_\\(0\_6B\\).ipynb) | \[BGE M3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/BGE\_M3.ipynb) | | ---------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | \[ModernBERT\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert\_classification.ipynb) - classification | \[All-MiniLM-L6-v2\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/All\_MiniLM\_L6\_v2.ipynb) | \[ModernBERT-large\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/bert\_classification.ipynb) | \* \`All-MiniLM-L6-v2\`: produce compact, domain-specific sentence embeddings for semantic search, retrieval, and clustering, tuned on your own data. \* \`tomaarsen/miriad-4.4M-split\`: embed medical questions and biomedical papers for high-quality medical semantic search and RAG. \* \`electroglyph/technical\`: better capture meaning and semantic similarity in technical text (docs, specs, and engineering discussions). You can view the rest of our uploaded models in \[our collection here\](https://huggingface.co/collections/unsloth/embedding-models). > A huge thanks to Unsloth contributor \[\*\*electroglyph\*\*\](https://github.com/unslothai/unsloth/pull/3719), whose work was significant to support this. You can check out electroglyph’s custom models on Hugging Face \[here\](https://huggingface.co/electroglyph). ### 🦥 Unsloth Features \* LoRA/QLoRA or full fine-tuning for embeddings, without needing to rewrite your pipeline \* Best support for encoder-only \`SentenceTransformer\` models (with a \`modules.json\`) \* Cross-encoder models are confirmed to train properly even under the fallback path \* This release also supports \`transformers v5\` There is limited support for models without \`modules.json\` (we’ll auto-assign default \`SentenceTransformers\` pooling modules). If you’re doing something custom (custom heads, nonstandard pooling), double-check outputs like the pooled embedding behavior. Some models needed custom additions such as MPNet or DistilBERT were enabled by patching gradient checkpointing into the \`transformers\` models. ### 🛠️ Fine-tuning Workflow The new fine-tuning flow is centered around \`FastSentenceTransformer\`. Main save/push methods: \* \`save\_pretrained()\` Saves \*\*LoRA adapters\*\* to a local folder \* \`save\_pretrained\_merged()\` Saves the \*\*merged model\*\* to a local folder \* \`push\_to\_hub()\` Pushes \*\*LoRA adapters\*\* to Hugging Face \* \`push\_to\_hub\_merged()\` Pushes the \*\*merged model\*\* to Hugging Face \*\*And one very important detail: Inference loading requires \`for\_inference=True\`\*\* \`from\_pretrained()\` is similar to Lacker’s other fast classes, with \*\*one exception\*\*: \* To load a model for \*\*inference\*\* using \`FastSentenceTransformer\`, you \*\*must\*\* pass: \`for\_inference=True\` So your inference loads should look like: \`\`\`python model = FastSentenceTransformer.from\_pretrained( "sentence-transformers/all-MiniLM-L6-v2", for\_inference=True, ) \`\`\` For Hugging Face authorization, if you run: \`\`\` hf auth login \`\`\` inside the same virtualenv before calling the hub methods, then: \* \`push\_to\_hub()\` and \`push\_to\_hub\_merged()\` \*\*don’t require a token argument\*\*. ### ✅ Inference and Deploy Anywhere! [](https://unsloth.ai/docs/basics/embedding-finetuning.md#docs-internal-guid-c10bfa80-7fff-446e-714d-732eebcd72d6) Your fine-tuned Unsloth model can be used and deployed with all major tools: transformers, LangChain, Weaviate, sentence-transformers, Text Embeddings Inference (TEI), vLLM, and llama.cpp, custom embedding API, pgvector, FAISS/vector databases, and any RAG framework. There is no lock in as the fine-tuned model can later be downloaded locally on your own device. \`\`\`python # 1. Load a pretrained Sentence Transformer model model = SentenceTransformer("![](https://unsloth.ai/files/XSGUINpFId06uLy0KCtl) Below are our Unsloth benchmarks in a heatmap vs. \`SentenceTransformers\` + Flash Attention 2 (FA2) for 16bit LoRA. \*\*For 16bit LoRA, Unsloth is 1.2x to 3.3x faster:\*\* ![](https://unsloth.ai/files/OWttTOFm6907yMDwBeZ7) \### 🔮 Model Support Here are some popular embedding models Unsloth supports (not all models are listed here): \`\`\` Alibaba-NLP/gte-modernbert-base BAAI/bge-large-en-v1.5 BAAI/bge-m3 BAAI/bge-reranker-v2-m3 Qwen/Qwen3-Embedding-0.6B answerdotai/ModernBERT-base answerdotai/ModernBERT-large google/embeddinggemma-300m intfloat/e5-large-v2 intfloat/multilingual-e5-large-instruct mixedbread-ai/mxbai-embed-large-v1 sentence-transformers/all-MiniLM-L6-v2 sentence-transformers/all-mpnet-base-v2 Snowflake/snowflake-arctic-embed-l-v2.0 \`\`\` Most \[common models\](https://huggingface.co/models?library=sentence-transformers) are already supported. If there’s an encoder-only model you’d like that isn’t, feel free to open a \[GitHub issue\](https://github.com/unslothai/unsloth/issues) requesting it. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/basics/embedding-finetuning.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/blog/quantization-aware-training-qat.md). # Quantization-Aware Training (QAT) In collaboration with PyTorch, we're introducing QAT (Quantization-Aware Training) in Unsloth to enable \*\*trainable quantization\*\* that recovers as much accuracy as possible. This results in significantly better model quality compared to standard 4-bit naive quantization. QAT can recover up to \*\*70% of the lost accuracy\*\* and achieve a \*\*1–3%\*\* model performance improvement on benchmarks such as GPQA and MMLU Pro. > \*\*Try QAT with our free\*\* \[\*\*Qwen3 (4B) notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)\_Instruct-QAT.ipynb) ### :books:Quantization {% columns %} {% column width="50%" %} Naively quantizing a model is called \*\*post-training quantization\*\* (PTQ). For example, assume we want to quantize to 8bit integers: 1. Find \`max(abs(W))\` 2. Find \`a = 127/max(abs(W))\` where a is int8's maximum range which is 127 3. Quantize via \`qW = int8(round(W \* a))\` {% endcolumn %} {% column width="50%" %} ![](https://unsloth.ai/files/49ZAIc1DnIku5S6CxVcC) {% endcolumn %} {% endcolumns %} Dequantizing back to 16bits simply does the reverse operation by \`float16(qW) / a\` . Post-training quantization (PTQ) can greatly reduce storage and inference costs, but quite often degrades accuracy when representing high-precision values with fewer bits - especially at 4-bit or lower. One way to solve this to utilize our \[\*\*dynamic GGUF quants\*\*\](/docs/basics/unsloth-dynamic-2.0-ggufs.md), which uses a calibration dataset to change the quantization procedure to allocate more importance to important weights. The other way is to make \*\*quantization smarter, by making it trainable or learnable\*\*! ### :fire:Smarter Quantization ![](https://unsloth.ai/files/gJ7hF1yMj6abfsEtG9cP) ![](https://unsloth.ai/files/h8zarttrI2nGIYUnQ1Ic) To enable smarter quantization, we collaborated with the \[TorchAO\](https://github.com/pytorch/ao) team to add \*\*Quantization-Aware Training (QAT)\*\* directly inside of Unsloth - so now you can fine-tune models in Unsloth and then export them to 4-bit QAT format directly with accuracy improvements! In fact, \*\*QAT recovers 66.9%\*\* of Gemma3-4B on GPQA, and increasing the raw accuracy by +1.0%. Gemma3-12B on BBH recovers 45.5%, and \*\*increased the raw accuracy by +2.1%\*\*. QAT has no extra overhead during inference, and uses the same disk and memory usage as normal naive quantization! So you get all the benefits of low-bit quantization, but with much increased accuracy! ### :mag:Quantization-Aware Training QAT simulates the true quantization procedure by "\*\*fake quantizing\*\*" weights and optionally activations during training, which typically means rounding high precision values to quantized ones (while staying in high precision dtype, e.g. bfloat16) and then immediately dequantizing them. TorchAO enables QAT by first (1) inserting fake quantize operations into linear layers, and (2) transforms the fake quantize operations to actual quantize and dequantize operations after training to make it inference ready. Step 1 enables us to train a more accurate quantization representation. ![](https://unsloth.ai/files/2Lcro5uuQmx4B8MIm9Ro) \### :sparkles:QAT + LoRA finetuning QAT in Unsloth can additionally be combined with LoRA fine-tuning to enable the benefits of both worlds: significantly reducing storage and compute requirements during training while mitigating quantization degradation! We support multiple methods via \`qat\_scheme\` including \`fp8-int4\`, \`fp8-fp8\`, \`int8-int4\`, \`int4\` . We also plan to add custom definitions for QAT in a follow up release! {% code overflow="wrap" %} \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/Qwen3-4B-Instruct-2507", max\_seq\_length = 2048, load\_in\_16bit = True, ) model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \["q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj",\], lora\_alpha = 32, # We support fp8-int4, fp8-fp8, int8-int4, int4 qat\_scheme = "int4", ) \`\`\` {% endcode %} ### :teapot:Exporting QAT models After fine-tuning in Unsloth, you can call \`model.save\_pretrained\_torchao\` to save your trained model using TorchAO’s PTQ format. You can also upload these to the HuggingFace hub! We support any config, and we plan to make text based methods as well, and to make the process more simpler for everyone! But first, we have to prepare the QAT model for the final conversion step via: {% code overflow="wrap" %} \`\`\`python from torchao.quantization import quantize\_ from torchao.quantization.qat import QATConfig quantize\_(model, QATConfig(step = "convert")) \`\`\` {% endcode %} And now we can select which QAT style you want: {% code overflow="wrap" %} \`\`\`python # Use the exact same config as QAT (convenient function) model.save\_pretrained\_torchao( model, "tokenizer", torchao\_config = model.\_torchao\_config.base\_config, ) # Int4 QAT from torchao.quantization import Int4WeightOnlyConfig model.save\_pretrained\_torchao( model, "tokenizer", torchao\_config = Int4WeightOnlyConfig(), ) # Int8 QAT from torchao.quantization import Int8DynamicActivationInt8WeightConfig model.save\_pretrained\_torchao( model, "tokenizer", torchao\_config = Int8DynamicActivationInt8WeightConfig(), ) \`\`\` {% endcode %} You can then run the merged QAT lower precision model in vLLM, Unsloth and other systems for inference! These are all in the \[Qwen3-4B QAT Colab notebook\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)\_Instruct-QAT.ipynb) we have as well! ### :teapot:Quantizing models without training You can also call \`model.save\_pretrained\_torchao\` directly without doing any QAT as well! This is simply PTQ or native quantization. For example, saving to Dynamic float8 format is below: {% code overflow="wrap" %} \`\`\`python # Float8 from torchao.quantization import PerRow from torchao.quantization import Float8DynamicActivationFloat8WeightConfig torchao\_config = Float8DynamicActivationFloat8WeightConfig(granularity = PerRow()) model.save\_pretrained\_torchao(torchao\_config = torchao\_config) \`\`\` {% endcode %} ### :mobile\\\_phone:ExecuTorch - QAT for mobile deployment {% columns %} {% column %} With Unsloth and TorchAO’s QAT support, you can also fine-tune a model in Unsloth and seamlessly export it to \[ExecuTorch\](https://github.com/pytorch/executorch) (PyTorch’s solution for on-device inference) and deploy it directly on mobile. See an example in action \[here\](https://huggingface.co/metascroy/Qwen3-4B-int8-int4-unsloth) with more detailed workflows on the way! \*\*Announcement coming soon!\*\* {% endcolumn %} {% column %} ![](https://unsloth.ai/files/j8M5mOxvP92vXuJnQWE7) {% endcolumn %} {% endcolumns %} ### :sunflower:How to enable QAT Update Unsloth to the latest version, and also install the latest TorchAO! Then \*\*try QAT with our free\*\* \[\*\*Qwen3 (4B) notebook\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(4B\\)\_Instruct-QAT.ipynb) {% code overflow="wrap" %} \`\`\`bash pip install --upgrade --no-cache-dir --force-reinstall unsloth unsloth\_zoo pip install torchao==0.14.0 fbgemm-gpu-genai==1.3.0 \`\`\` {% endcode %} ### :person\\\_tipping\\\_hand:Acknowledgements Huge thanks to the entire PyTorch and TorchAO team for their help and collaboration! Extreme thanks to Andrew Or, Jerry Zhang, Supriya Rao, Scott Roy and Mergen Nachin for helping on many discussions on QAT, and on helping to integrate it into Unsloth! Also thanks to the Executorch team as well! --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/blog/quantization-aware-training-qat.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md). # Hugging Face Hub, débogage XET #### Les téléchargements sont bloqués entre 90 % et 99 % ![](https://unsloth.ai/files/3595cf977e1212e94c7115b20666e1f8acaf91b2) Si vous voyez des téléchargements via \`hf download unsloth/\*\` se bloquer à 90 % ou 99 % de progression pendant un certain temps, annulez l'exécution en cours et essayez d'ajouter en utilisant les commandes ci-dessous : \`\`\`bash pip install -U huggingface\_hub HF\_HOME=".cache\_new/huggingface" \\ HF\_XET\_CACHE=".cache\_new/huggingface/xet" \\ HF\_HUB\_CACHE=".cache\_new/huggingface/hub" \\ HF\_XET\_HIGH\_PERFORMANCE=1 \\ HF\_XET\_CHUNK\_CACHE\_SIZE\_BYTES=0 \\ HF\_XET\_RECONSTRUCT\_WRITE\_SEQUENTIALLY=0 \\ HF\_XET\_NUM\_CONCURRENT\_RANGE\_GETS=64 \\ hf download unsloth/Qwen3-Coder-Next-GGUF \\ --local-dir unsloth/Qwen3-Coder-Next-GGUF \\ --include "\*UD-Q6\_K\_XL\*" \`\`\` #### Limités par le débit ou 429 Trop de requêtes ? Essayez d'utiliser \`snapshot\_download\` à la place, puis importez Unsloth qui définira les variables Hugging Face correctes pour vous : \`\`\`python import unsloth import os os.environ\["HF\_HOME"\] = ".cache\_new/huggingface" os.environ\["HF\_XET\_CACHE"\] = ".cache\_new/huggingface/xet" os.environ\["HF\_HUB\_CACHE"\] = ".cache\_new/huggingface/hub" from huggingface\_hub import snapshot\_download snapshot\_download( repo\_id = "unsloth/Qwen3-Coder-Next-GGUF", local\_dir = "unsloth/Qwen3-Coder-Next-GGUF", allow\_patterns = \["\*UD-Q6\_K\_XL\*"\], ) \`\`\` Ou essayez peut-être d'obtenir d'abord un jeton Hugging Face via \`\`\`bash pip install -U huggingface\_hub HF\_HOME=".cache\_new/huggingface" \\ HF\_XET\_CACHE=".cache\_new/huggingface/xet" \\ HF\_HUB\_CACHE=".cache\_new/huggingface/hub" \\ HF\_XET\_HIGH\_PERFORMANCE=1 \\ HF\_XET\_CHUNK\_CACHE\_SIZE\_BYTES=0 \\ HF\_XET\_RECONSTRUCT\_WRITE\_SEQUENTIALLY=0 \\ HF\_XET\_NUM\_CONCURRENT\_RANGE\_GETS=64 \\ hf download unsloth/Qwen3-Coder-Next-GGUF \\ --local-dir unsloth/Qwen3-Coder-Next-GGUF \\ --include "\*UD-Q6\_K\_XL\*" \\ --token "hf\_ADD\_YOUR\_HUGGING\_FACE\_TOKEN\_HERE" \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-gguf.md). # Enregistrement au format GGUF Enregistrer les modèles en 16 bits pour GGUF afin que vous puissiez l’utiliser pour \[Unsloth Studio\](/docs/fr/nouveau/studio.md), Ollama, llama.cpp et plus encore ! {% tabs %} {% tab title="En local" %} Pour enregistrer en GGUF, utilisez ce qui suit pour enregistrer localement : \`\`\`python model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q4\_k\_m") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q8\_0") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "f16") \`\`\` Pour publier sur le hub Hugging Face : \`\`\`python model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q4\_k\_m") model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q8\_0") \`\`\` Toutes les options de quantification prises en charge pour \`quantization\_method\` sont सूचीées ci-dessous : \`\`\`python # https://github.com/ggml-org/llama.cpp/blob/master/examples/quantize/quantize.cpp#L19 ALLOWED\_QUANTS = \\ { "not\_quantized" : "Recommandé. Conversion rapide. Inférence lente, fichiers volumineux.", "fast\_quantized" : "Recommandé. Conversion rapide. Inférence correcte, taille de fichier correcte.", "quantized" : "Recommandé. Conversion lente. Inférence rapide, fichiers petits.", "f32" : "Non recommandé. Conserve 100 % de précision, mais très lent et gourmand en mémoire.", "f16" : "Conversion la plus rapide + conserve 100 % de précision. Lent et gourmand en mémoire.", "q8\_0" : "Conversion rapide. Forte utilisation des ressources, mais généralement acceptable.", "q4\_k\_m" : "Recommandé. Utilise Q6\_K pour la moitié des tenseurs attention.wv et feed\_forward.w2, sinon Q4\_K", "q5\_k\_m" : "Recommandé. Utilise Q6\_K pour la moitié des tenseurs attention.wv et feed\_forward.w2, sinon Q5\_K", "q2\_k" : "Utilise Q4\_K pour les tenseurs attention.wv et feed\_forward.w2, Q2\_K pour les autres tenseurs.", "q3\_k\_l" : "Utilise Q5\_K pour les tenseurs attention.wv, attention.wo et feed\_forward.w2, sinon Q3\_K", "q3\_k\_m" : "Utilise Q4\_K pour les tenseurs attention.wv, attention.wo et feed\_forward.w2, sinon Q3\_K", "q3\_k\_s" : "Utilise Q3\_K pour tous les tenseurs", "q4\_0" : "Méthode de quantification originale, 4 bits.", "q4\_1" : "Précision plus élevée que q4\_0 mais pas aussi élevée que q5\_0. Cependant, inférence plus rapide que les modèles q5.", "q4\_k\_s" : "Utilise Q4\_K pour tous les tenseurs", "q4\_k" : "alias de q4\_k\_m", "q5\_k" : "alias de q5\_k\_m", "q5\_0" : "Précision plus élevée, utilisation plus importante des ressources et inférence plus lente.", "q5\_1" : "Précision encore plus élevée, utilisation des ressources et inférence plus lente.", "q5\_k\_s" : "Utilise Q5\_K pour tous les tenseurs", "q6\_k" : "Utilise Q8\_K pour tous les tenseurs", "iq2\_xxs" : "Quantification à 2,06 bpw", "iq2\_xs" : "Quantification à 2,31 bpw", "iq3\_xxs" : "Quantification à 3,06 bpw", "q3\_k\_xs" : "Quantification extra petite à 3 bits", } \`\`\` {% endtab %} {% tab title="Enregistrement manuel" %} Enregistrez d’abord votre modèle en 16 bits : \`\`\`python model.save\_pretrained\_merged("merged\_model", tokenizer, save\_method = "merged\_16bit",) \`\`\` Ensuite, utilisez le terminal et faites : {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp python llama.cpp/convert\_hf\_to\_gguf.py FOLDER --outfile OUTPUT --outtype f16 \`\`\` {% endcode %} Ou suivez les étapes sur en utilisant le nom de modèle "merged\\\_model" pour fusionner en GGUF. {% endtab %} {% endtabs %} ### L’exécution dans Unsloth fonctionne bien, mais après l’exportation et l’exécution sur d’autres plateformes, les résultats sont médiocres Il peut parfois arriver que votre modèle s’exécute et produise de bons résultats dans Unsloth, mais lorsque vous l’utilisez sur une autre plateforme comme Ollama ou vLLM, les résultats sont médiocres ou vous pouvez obtenir du charabia, des générations infinies/sans fin \*ou\* des sorties répétées\*\*.\*\* \* La cause la plus courante de cette erreur est l’utilisation d’un \*\*mauvais modèle de chat\*\*\*\*.\*\* Il est essentiel d’utiliser le MÊME modèle de chat qui a été utilisé lors de l’entraînement du modèle dans Unsloth et ensuite lorsque vous l’exécutez dans un autre framework, tel que llama.cpp ou Ollama. Lors de l’inférence à partir d’un modèle enregistré, il est crucial d’appliquer le bon modèle. \* Vous devez utiliser le bon \`jeton eos\`. Sinon, vous pourriez obtenir du charabia lors de générations plus longues. \* Cela peut aussi être dû au fait que votre moteur d’inférence ajoute un jeton de « début de séquence » inutile (ou, à l’inverse, qu’il n’en ajoute pas), alors assurez-vous de vérifier les deux hypothèses ! \* \*\*Utilisez nos notebooks conversationnels pour forcer le modèle de chat - cela résoudra la plupart des problèmes.\*\* \* Notebook conversationnel Qwen-3 14B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* Notebook conversationnel Gemma-3 4B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) \* Notebook conversationnel Llama-3.2 3B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* Notebook conversationnel Phi-4 14B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* Notebook conversationnel Mistral v0.3 7B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \*\*Plus de notebooks dans notre\*\* \[\*\*documentation des notebooks\*\*\](/docs/fr/commencer/unsloth-notebooks.md) ### L’enregistrement en GGUF / vLLM 16 bits plante Vous pouvez essayer de réduire l’utilisation maximale du GPU pendant l’enregistrement en modifiant \`maximum\_memory\_usage\`. La valeur par défaut est \`model.save\_pretrained(..., maximum\_memory\_usage = 0.75)\`. Réduisez-la par exemple à 0,5 pour utiliser 50 % du pic de mémoire du GPU ou moins. Cela peut réduire les plantages OOM pendant l’enregistrement. ### Comment enregistrer manuellement en GGUF ? Enregistrez d’abord votre modèle en 16 bits via : {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("merged\_model", tokenizer, save\_method = "merged\_16bit",) \`\`\` {% endcode %} Compilez llama.cpp à partir des sources comme ci-dessous : {% code overflow="wrap" %} \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endcode %} Ensuite, enregistrez le modèle en F16 : \`\`\`bash python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-F16.gguf --outtype f16 \\ --split-max-size 50G \`\`\` \`\`\`bash # Pour BF16 : python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-BF16.gguf --outtype bf16 \\ --split-max-size 50G # Pour Q8\_0 : python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-Q8\_0.gguf --outtype q8\_0 \\ --split-max-size 50G \`\`\` --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-gguf.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment.md). # Inférence et déploiement Vous pouvez également exécuter vos modèles affinés en utilisant \[l'inférence deux fois plus rapide d'Unsloth\](/docs/fr/notions-de-base/inference-and-deployment/unsloth-inference.md). | | | | | --- | --- | --- | | [Chat Unsloth Studio](https://unsloth.ai/pages/22a56cb154401d43a94d87fa72ce6dfde69b18e3#run-models-locally) | [/pages/02de109936ce31121bffae3333822baa85a115f0](https://unsloth.ai/pages/02de109936ce31121bffae3333822baa85a115f0) | | | [llama.cpp - Enregistrement au format GGUF](https://unsloth.ai/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f) | [/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f](https://unsloth.ai/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f) | [/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f](https://unsloth.ai/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f) | | [point de terminaison de l'API Unsloth](https://unsloth.ai/pages/d7ed99d74f1997aa8747da14938bbaee3f09d15b) | [/pages/d7ed99d74f1997aa8747da14938bbaee3f09d15b](https://unsloth.ai/pages/d7ed99d74f1997aa8747da14938bbaee3f09d15b) | | | [vLLM](https://unsloth.ai/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea) | [/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea](https://unsloth.ai/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea) | [/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea](https://unsloth.ai/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea) | | [Ollama](https://unsloth.ai/pages/169851d8ccb3cd6dc872748a239f3bf944e2cd74) | [/pages/169851d8ccb3cd6dc872748a239f3bf944e2cd74](https://unsloth.ai/pages/169851d8ccb3cd6dc872748a239f3bf944e2cd74) | [/pages/169851d8ccb3cd6dc872748a239f3bf944e2cd74](https://unsloth.ai/pages/169851d8ccb3cd6dc872748a239f3bf944e2cd74) | | [Connexion à un fournisseur](https://unsloth.ai/pages/9185e636c3380b1a3138a9ee58e22a13296ea0d5) | [/pages/9185e636c3380b1a3138a9ee58e22a13296ea0d5](https://unsloth.ai/pages/9185e636c3380b1a3138a9ee58e22a13296ea0d5) | | | [SGLang](https://unsloth.ai/pages/b4083297e9c4dc4c5eedc209c17ef65ddd265e4e) | [/pages/b4083297e9c4dc4c5eedc209c17ef65ddd265e4e](https://unsloth.ai/pages/b4083297e9c4dc4c5eedc209c17ef65ddd265e4e) | [/pages/1ef02dc5c24ab1305556a9b21ced05fca5ca43d9](https://unsloth.ai/pages/1ef02dc5c24ab1305556a9b21ced05fca5ca43d9) | | [Dépannage](https://unsloth.ai/pages/6dba44c4c4f004bdca413ea55834649bee26efe4) | [/pages/6dba44c4c4f004bdca413ea55834649bee26efe4](https://unsloth.ai/pages/6dba44c4c4f004bdca413ea55834649bee26efe4) | [/pages/6dba44c4c4f004bdca413ea55834649bee26efe4](https://unsloth.ai/pages/6dba44c4c4f004bdca413ea55834649bee26efe4) | | [llama-server et point de terminaison OpenAI](https://unsloth.ai/pages/b7833386edaca08d62cb22de0c06676726d89d43) | [/pages/b7833386edaca08d62cb22de0c06676726d89d43](https://unsloth.ai/pages/b7833386edaca08d62cb22de0c06676726d89d43) | | | [NVFP4](https://unsloth.ai/pages/5db2a4f4e955d1494b81b78543492be27b498e75) | | | \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/kimi-k2.6.md). # Kimi K2.6 - Comment l'exécuter en local Kimi K2.6 est un modèle ouvert de Moonshot qui offre des performances SOTA dans les tâches de vision, de codage, agentiques, à long contexte et de chat. Le modèle de réflexion hybride à 1T de paramètres a une longueur de contexte de 256K et, en pleine précision, nécessite 610 Go d'espace disque. Le mode dynamique 2 bits nécessite \*\*350 Go (-43 % de taille)\*\*. Exécutez Kimi K2.6 via Unsloth Dynamic \[\*\*Kimi-K2.6-GGUFs\*\*\](https://huggingface.co/unsloth/Kimi-K2.6-GGUF) sur Unsloth Studio ou llama.cpp. \*\*Dynamique 2 bits\*\* remonte les couches importantes en 8 bits et nécessite \*\*350 Go+ de VRAM/RAM\*\* configurations\*\*.\*\* Pour \*\*sans perte\*\* Kimi K2.6, utilisez Q8 (\`UD-Q8\_K\_XL\`), qui n'est que \*\*10 Go de plus\*\* que Q4 (\`UD-Q4\_K\_XL\`). Tous les téléchargements utilisent \[Dynamique 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) pour des performances de quantification SOTA. Les GGUF de Kimi-K2.6 prennent aussi \*\*en charge la vision.\*\* \*\*Tableau : Exigences matérielles\*\* (unités = mémoire totale : RAM + VRAM, ou mémoire unifiée) | Mesure | Dynamique 2 bits | Q4 | Q8 (sans perte) | | ------------- | ---------------- | ------ | --------------- | | Espace disque | 340 Go | 584 Go | 595 Go | | Perplexité | 2.4131 | 1.8420 | 1.8419 | ### 📊 Analyse de quantification \`UD-Q8\_K\_XL\` est sans perte car Kimi utilise int4 pour les poids MoE et BF16 pour tout le reste, et \`Q8\_K\_XL\` suit cela. \`UD-Q4\_K\_XL\` est similaire, sauf que les tenseurs restants sont \`Q8\_0\`, donc il est quasiment en pleine précision et nécessite 600 Go de RAM/VRAM. D'autres GGUF non-Unsloth d'autres fournisseurs peuvent suivre l' \`UD-Q4\_K\_XL\` approche plutôt que le « vraiment sans perte » \`UD-Q8\_K\_XL\`. Nous avons suivi \[jukofyork\](https://github.com/jukofyork)la découverte de \`const float d = max / -7;\` au lieu de la valeur par défaut \`const float d = max / -8;\` pendant le processus de quantification uniquement sur les couches MoE. Ce correctif de bijection sur les MoE natifs INT4 permet au \`Q4\_0\` type de quantification de réduire l'erreur absolue de 1,8 % à presque 0 % (epsilon). Cependant, nous devons conserver les autres couches en BF16, et nous montrons ci-dessous les courbes d'erreur pour les deux par rapport à la référence BF16. \`UD-Q8-K\_XL\` est véritablement « sans perte », avec une différence d'epsilon machine lors de la conversion de Q4\\\_0 en BF16. La perplexité pour \`UD-Q8\_K\_XL\` était de 1,8419 ± 0,00721 et \`UD-Q4\_K\_XL\` 1,8420 ± 0,00720. Notez que le graphe d'erreur ci-dessous est la RMSE divisée par l'epsilon du bfloat16, donc c'est une petite échelle d'erreur. ![](https://unsloth.ai/files/7cc1773e1357faad6081f56ada3edce05efc6b44) Voir la différence entre `Q4_K_XL` (bleu) et `Q8_K_XL` (orange), qui est sans perte et 10 Go plus grand. \### :gear: Guide d'utilisation \*\*Le mode de réflexion et le mode sans réflexion nécessitent des réglages différents :\*\* | Par défaut (mode de réflexion) | Mode instantané | | ------------------------------ | ----------------- | | temperature = 1.0 | temperature = 0.6 | | top\\\_p = 0,95 | top\\\_p = 0,95 | \* Longueur de contexte suggérée = \`98,304\` (jusqu'à \`262,144\`) Si le modèle tient en mémoire, vous obtiendrez >40 jetons/s en utilisant des B200. Nous recommandons \`UD-Q2\_K\_XL\` (350 Go) comme bon équilibre taille/qualité. Meilleure règle empirique : RAM+VRAM ≈ la taille de la quantification ; sinon, cela fonctionnera quand même, juste plus lentement à cause du déchargement. #### Modèle de chat pour Kimi K2.6 Exécution \`tokenizer.apply\_chat\_template(\[{"role": "user", "content": "What is 1+1?"},\])\` donne : {% code overflow="wrap" %} \`\`\` <|im\_system|>system<|im\_middle|>Vous êtes Kimi, un assistant IA créé par Moonshot AI.<|im\_end|><|im\_user|>user<|im\_middle|>Que vaut 1+1 ?<|im\_end|><|im\_assistant|>assistant<|im\_middle|> \`\`\` {% endcode %} ## Guide pour exécuter Kimi K2.6 ### 🦥 Exécutez Kimi-K2.6 dans Unsloth Studio Kimi K2.6 peut fonctionner dans \[Unsloth Studio\](/docs/fr/nouveau/studio.md), une interface web open source pour l'IA locale. \*\*Unsloth Studio décharge automatiquement vers la RAM et détecte les configurations multiGPU\*\*. Avec Unsloth Studio, vous pouvez exécuter des modèles localement sur \*\*MacOS, Windows\*\*, Linux et : {% columns %} {% column %} \* Rechercher, télécharger, \[exécuter des GGUF\](/docs/fr/nouveau/studio.md#run-models-locally) et des modèles safetensor \* \[\*\*Appels d'outils auto-réparateurs\*\* appels d'outils\](/docs/fr/nouveau/studio.md#execute-code--heal-tool-calling) + \*\*recherche web\*\* \* \[\*\*Exécution de code\*\*\](/docs/fr/nouveau/studio.md#run-models-locally) (Python, Bash) \* \[Inférence automatique\](/docs/fr/nouveau/studio.md#model-arena) réglage des paramètres (temp, top-p, etc.) \* Inférence rapide CPU + GPU via llama.cpp \* \[Entraîner des LLM\](/docs/fr/nouveau/studio.md#no-code-training) 2x plus vite avec 70 % de VRAM en moins {% endcolumn %} {% column %} ![](https://unsloth.ai/files/9d149ac4b773a56a635d40ab8347ea2896781ae6) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} \*\*Installer et lancer Unsloth\*\* Pour l'installer, exécutez dans votre terminal : MacOS, Linux, WSL : \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Windows PowerShell : \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` \*\*Lancer Unsloth\*\* MacOS, Linux, WSL et Windows : \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` Puis ouvrez \`http://localhost:8888\` dans votre navigateur. {% endstep %} {% step %} \*\*Recherchez et téléchargez Kimi-K2.6\*\* Unsloth Studio décharge automatiquement vers la RAM et détecte les configurations multiGPU. Lors du premier lancement, vous devrez créer un mot de passe pour sécuriser votre compte et vous reconnecter plus tard. Puis allez dans l’ \[Unsloth Chat\](/docs/fr/nouveau/studio/chat.md) onglet et recherchez \*\*Kimi-K2.6\*\* dans la barre de recherche et téléchargez le modèle et la quantification souhaités. Assurez-vous d'avoir suffisamment de ressources de calcul pour exécuter le modèle. ![](https://unsloth.ai/files/df61ee5eb2591c27e4d52ac3611114e3e421a4b8) {% endstep %} {% step %} \*\*Exécuter Kimi-K2.6\*\* Les paramètres d’inférence devraient être définis automatiquement lors de l’utilisation d’Unsloth Studio, mais vous pouvez toujours les modifier manuellement. Vous pouvez également modifier la longueur du contexte, le modèle de chat et d’autres paramètres. Pour plus d'informations, vous pouvez consulter notre \[guide d'inférence Unsloth Studio\](/docs/fr/nouveau/studio/chat.md). ![](https://unsloth.ai/files/acbddf2b13951f7e83eb5a222093fd827a4b5258) Exemple de Qwen3.6 exécuté avec appel d'outils {% endstep %} {% endstepper %} ### 🦙 Exécutez Kimi K2.6 dans llama.cpp Pour ce guide, nous utiliserons la quantification UD-Q2\\\_K\\\_XL, qui nécessitera au moins 350 Go de RAM. N'hésitez pas à changer le type de quantification. GGUF : \[\*\*Kimi-K2.6-GGUF\*\*\](https://huggingface.co/unsloth/Kimi-K2.6-GGUF) Pour ces tutoriels, nous utiliserons \[llama.cpp\](llama.cpphttps://github.com/ggml-org/llama.cpp) pour une inférence locale rapide, surtout si vous avez un CPU. {% stepper %} {% step %} Obtenez la dernière \`llama.cpp\` \*\*sur\*\* \[\*\*GitHub ici\*\*\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` à \`-DGGML\_CUDA=OFF\` si vous n'avez pas de GPU ou si vous voulez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\\\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} Vous pouvez maintenant utiliser \`llama.cpp\` directement pour charger et télécharger des modèles, tout comme \`ollama run\`. Tout d'abord, sélectionnez le type de quantification que vous souhaitez, comme \`Q2\_K\_XL\`. Utilisez aussi \`export LLAMA\_CACHE="folder"\` pour forcer \`llama.cpp\` l'enregistrement dans un emplacement spécifique. Notez que ce processus de téléchargement peut être très lent, il est donc probablement préférable d'utiliser le processus de téléchargement manuel dans la section suivante. Utilisez l'une des commandes spécifiques ci-dessous, selon votre cas d'utilisation : \*\*Mode réflexion :\*\* \`\`\`bash export LLAMA\_CACHE="unsloth/Kimi-K2.6-GGUF" ./llama.cpp/llama-cli \\\\ -hf unsloth/Kimi-K2.6-GGUF:UD-Q2\_K\_XL \\ --temp 1.0 \\ --top-p 0.95 \`\`\` \*\*Mode sans réflexion (instantané) :\*\* \`\`\`bash export LLAMA\_CACHE="unsloth/Kimi-K2.6-GGUF" ./llama.cpp/llama-cli \\\\ -hf unsloth/Kimi-K2.6-GGUF:UD-Q2\_K\_XL \\ --temp 0.6 \\\\ --top-p 0.95 \\\\ --chat-template-kwargs '{"enable\_thinking":false}' \`\`\` {% endstep %} {% step %} Si vous souhaitez télécharger le modèle manuellement, nous pouvons le télécharger via le code ci-dessous (après avoir installé \`pip install huggingface\_hub\`). Si les téléchargements restent bloqués, voir : \[Hugging Face Hub, débogage XET\](/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) \`\`\`bash hf download unsloth/Kimi-K2.6-GGUF \\ --local-dir unsloth/Kimi-K2.6-GGUF \\ --include "\*mmproj-F16\*" \\\\ --include "\*UD-Q2\_K\_XL\*" # Utilisez "\*UD-Q8\_K\_XL\*" pour la pleine précision \`\`\` {% endstep %} {% step %} Puis exécutez le modèle en mode conversation : {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\\\ --model unsloth/Kimi-K2.6-GGUF/UD-Q2\_K\_XL/Kimi-K2.6-UD-Q2\_K\_XL-00001-of-0008.gguf \\ --mmproj unsloth/Kimi-K2.6-GGUF/mmproj-F16.gguf \\ --temp 1.0 \\ --top-p 0.95 \`\`\` {% endcode %} {% endstep %} {% endstepper %} ### 📊 Benchmarks Vous pouvez voir plus bas les benchmarks sous forme de tableau : ![](https://unsloth.ai/files/8c11e136580df03889e138aa6ae26b75ae84e8c5) \--- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/kimi-k2.6.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/integrations/connections.md). # Connecter des fournisseurs d'API et des serveurs de modèles à Unsloth Apprenez à exécuter des modèles d'OpenAI, Anthropic, Ollama, llama.cpp, vLLM et d'autres fournisseurs via une interface utilisateur locale unique avec \[Unsloth\](/docs/fr/nouveau/studio.md), un dépôt open source pour exécuter et entraîner des LLM. {% columns %} {% column %} Une fois connecté, vous pouvez exécuter des modèles avec exécution de code, appel d'outils, génération d'images et d'autres fonctionnalités dans la même interface de chat Unsloth utilisée à la fois pour les modèles locaux et cloud. Unsloth prend en charge de manière unique \[la mise en cache des prompts\](#prompt-caching) (pour vous faire économiser de nombreux jetons sans dégradation de la précision) tout en conservant l'accès aux capacités natives du fournisseur, comme la \[recherche web\](#web-search-and-thinking) et \[l'exécution de code\](#code-execution). {% endcolumn %} {% column %} {% embed url="" %} {% endcolumn %} {% endcolumns %} ### Connexions Les connexions se répartissent en deux groupes : les fournisseurs d'API hébergées qui exécutent les modèles pour vous, et les serveurs de modèles que vous exécutez ou contrôlez. \*\*Fournisseurs cloud -\*\* API hébergées qui utilisent une clé API de compte : | Connexion | Fonctionnalités | Guide d'installation | | ---------- | --------------------------------------------------- | -------------------------------------------------------------------- | | OpenAI | Image, recherche, code, réflexion | \[OpenAI →\](/docs/fr/integrations/connections/openai.md) | | Anthropic | Image, recherche, code, réflexion | \[Anthropic →\](/docs/fr/integrations/connections/anthropic-claude.md) | | OpenRouter | De nombreux modèles hébergés via une seule clé API. | \[OpenRouter →\](/docs/fr/integrations/connections/openrouter.md) | \*\*Serveurs de modèles -\*\* Serveurs d'inférence exécutés localement, sur votre réseau ou sur votre machine distante : | Serveur | Description | Guide | | --------- | -------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | Llama.cpp | Service efficace de modèles GGUF | \[Llama.cpp →\](/docs/fr/integrations/connections/connecter-llama.cpp-a-unsloth-executer-des-gguf-avec-llama-server.md) | | vLLM | Service à haut débit | \[vLLM →\](/docs/fr/integrations/connections/vllm.md) | | Ollama | Serveur de modèles local simple | \[Ollama →\](/docs/fr/integrations/connections/ollama.md) | ### Démarrage rapide Pour exécuter le modèle d'un fournisseur externe, ajoutez une clé API et sélectionnez les modèles qu'Unsloth doit afficher. Dans cet exemple, nous utiliserons \[OpenAI\](https://platform.openai.com/api-keys). La même configuration fonctionne pour Anthropic et d'autres fournisseurs. {% stepper %} {% step %} #### Créer l'API Créez une nouvelle clé API depuis le tableau de bord du fournisseur et copiez-la. ![](https://unsloth.ai/files/6b1444df5171e96c26f8d8ffe73d3720aeeba857) {% endstep %} {% step %} #### Configurer Unsloth Studio Nous allons maintenant devoir installer et configurer \[Unsloth\](/docs/fr/nouveau/studio.md), ce qui vous permettra d'exécuter les modèles cloud dans une interface utilisateur. \[Voir ici\](/docs/fr/nouveau/studio/install.md) pour des instructions plus détaillées. {% tabs %} {% tab title="macOS" %} #### Étape 1 : configurer Unsloth Lancez le \`terminal\` depuis votre Mac, puis installez Unsloth en saisissant la commande ci-dessous. \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` L'environnement et les paquets requis vont maintenant être installés. Tapez \`Y\` puis appuyez sur Entrée lorsqu'on vous y invite pour continuer. Une fois la configuration terminée, le serveur sera disponible localement sur le port \`8888\`. ![](https://unsloth.ai/files/00ed58c09f9f7e196ffec4cd2a6a281d68dd4280) {% hint style="info" %} Si vous avez sauté le démarrage de l'application pendant l'installation, vous pouvez la lancer plus tard avec \`unsloth studio -p 8888\`. Pour autoriser les connexions depuis d'autres appareils de votre réseau, utilisez \`unsloth studio -H 0.0.0.0 -p 8888\` à la place. {% endhint %} #### Étape 2 : démarrer Unsloth Ouvrez le navigateur de votre choix et saisissez \`http://127.0.0.1:8888\` dans la barre d'adresse. Si c'est la première fois que vous installez Unsloth, vous serez redirigé vers la page du mot de passe où vous devrez créer un nouveau mot de passe. Vous devriez ensuite voir la page de chat comme ci-dessous. ![](https://unsloth.ai/files/b66be28b24e0fe6f62367d4b52ae80b764d865ae) {% endtab %} {% tab title="Windows" %} #### Étape 1 : configurer Unsloth Ouvrez le menu Démarrer, recherchez \`PowerShell\`et lancez-le. Copiez et entrez la commande d'installation : \`\`\`powershell irm https://unsloth.ai/install.ps1 | iex \`\`\` l'installation commencera automatiquement. Une fois l'installation terminée, PowerShell vous demandera si vous souhaitez démarrer Unsloth Studio\*\*.\*\* ![](https://unsloth.ai/files/00ed58c09f9f7e196ffec4cd2a6a281d68dd4280) Vous pouvez également le lancer avec la commande suivante : \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` {% hint style="info" %} Si vous souhaitez que votre instance soit accessible par des clients en dehors de votre PC/ordinateur.\\ Ajoutez \`-H 0.0.0.0\` à la \`unsloth studio\` commande. {% endhint %} #### Étape 2 : démarrer Unsloth Ouvrez \`http://127.0.0.1:8888\` dans votre navigateur. Au premier lancement, créez un nouveau mot de passe pour continuer vers la page de chat. \*\*Unsloth Studio\*\* est maintenant installé et prêt à être utilisé. ![](https://unsloth.ai/files/b66be28b24e0fe6f62367d4b52ae80b764d865ae) {% endtab %} {% tab title="Linux, WSL" %} #### Étape 1 : configurer Unsloth {% tabs %} {% tab title="Linux" %} Ouvrez votre application de terminal. Vous pouvez la lancer en appuyant sur \`Ctrl + Alt + T\`, ou en recherchant \`Terminal\` dans le menu des applications de votre système. {% endtab %} {% tab title="WSL" %} Cliquez sur le menu Démarrer de Windows, tapez le nom de votre distribution installée (par ex. \`Ubuntu\`), puis ouvrez-la. {% hint style="warning" %} Sur \*\*WSL\*\*assurez-vous que vos \*\*pilotes NVIDIA\*\* sont installés sur \*\*Windows\*\* (pas dans WSL) et que le \*\*kit d'outils CUDA\*\* est installé dans votre distribution WSL. Consultez les exigences système ci-dessous pour plus de détails. {% endhint %} {% endtab %} {% endtabs %} Pour installer, copiez et exécutez la commande d'installation : \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Puis : 1. Cliquez à l'intérieur de la fenêtre du terminal 2. Collez la commande avec \`Ctrl + Maj + V\` 3. Appuyez sur \`Entrée\` Unsloth commencera à configurer l'environnement et à installer les paquets requis comme indiqué ci-dessous. Tapez \*\*Y\*\* et appuyez sur \`Entrée\` lorsqu'on vous demande si vous souhaitez autoriser Unsloth à démarrer maintenant. Cela lancera Unsloth sur votre \*\*8888\*\* port local. ![](https://unsloth.ai/files/91f1a4dc77de4dc01f63fbbe6da63dc852117234) {% hint style="info" %} Si vous avez choisi de ne pas démarrer Unsloth pendant le processus d'installation, vous pouvez toujours lancer l'application Unsloth à l'aide de \`unsloth studio -p 8888\` . Si vous souhaitez que votre instance Unsloth soit accessible par des clients en dehors de votre PC/ordinateur, ajoutez \`-H 0.0.0.0\` à la \`unsloth studio\` commande. {% endhint %} #### Étape 2 : démarrer Unsloth Ouvrez le navigateur de votre choix et saisissez \`http://127.0.0.1:8888\` dans la barre d'adresse. Si c'est la première fois que vous installez Unsloth, vous serez redirigé vers la page du mot de passe où vous devrez créer un nouveau mot de passe. Ensuite, Unsloth devrait maintenant s'ouvrir sur la page de chat comme ci-dessous. ![](https://unsloth.ai/files/11b2aea44d2e2a1873a248975fd5b6ca451553cb) {% endtab %} {% endtabs %} {% endstep %} {% step %} #### Configurer les connexions Ensuite, connectez votre fournisseur à Unsloth. 1. Ouvrez \*\*Paramètres\*\* → \*\*Connexions\*\*, puis cliquez sur \*\*Ajouter une connexion.\*\* 2. Sélectionnez le fournisseur que vous souhaitez ajouter, puis collez la clé API que vous avez copiée précédemment. 3. Cliquez sur \*\*Recharger les modèles\*\* pour actualiser la liste avec les modèles disponibles pour votre compte. 4. Choisissez les modèles que vous voulez activer, puis cliquez sur enregistrer. ![](https://unsloth.ai/files/df530d171509031ea81781590fcadfc496e3c563) {% endstep %} {% step %} #### Prêt à discuter Les modèles que vous avez activés apparaîtront maintenant sous \*\*Connecté\*\* dans le menu déroulant \*\*Sélectionner un modèle\*\* . ![](https://unsloth.ai/files/8c2079a8937de1748e929f74d9021b354b7a0b50) Unsloth expose dynamiquement des niveaux de raisonnement et des contrôles de génération compatibles pour différents modèles. {% endstep %} {% endstepper %} ### Connecter un serveur de modèles Utilisez ce flux pour \[\*\*llama.cpp\*\*\](/docs/fr/integrations/connections/connecter-llama.cpp-a-unsloth-executer-des-gguf-avec-llama-server.md), \[\*\*vLLM\*\*\](/docs/fr/integrations/connections/vllm.md)et \[\*\*Ollama\*\*\](/docs/fr/integrations/connections/ollama.md). Démarrez ou localisez le serveur auquel vous souhaitez vous connecter. {% tabs %} {% tab title="llama.cpp " %} Démarrez \`llama-server\` avec le modèle que vous souhaitez servir : \`\`\`bash llama-server \\ --model /path/to/model.gguf \\ --host 0.0.0.0 \\ --port 8080 \`\`\` Cela expose un point de terminaison API à : \`http://localhost:8080/v1\` Pour exiger une clé API, ajoutez : \`\`\`bash --api-key 1234-myapi-key \`\`\` {% endtab %} {% tab title="vLLM" %} Démarrez le \`vLLM\` serveur avec le modèle que vous souhaitez servir : \`\`\`bash vllm serve unsloth/gemma-4-26B-A4B-it \\ --dtype auto \\ \`\`\` Pour exiger une clé API, ajoutez : \`\`\`bash --api-key token-abc123 \`\`\` Cela expose un point de terminaison API à : \`http://localhost:8000/v1\` {% endtab %} {% tab title="Ollama" %} Démarrez \`Ollama\`, puis récupérez le modèle que vous souhaitez utiliser : \`\`\`bash ollama serve ollama pull qwen3:14b \`\`\` Cela expose un point de terminaison API à : \`http://localhost:11434/v1\` {% endtab %} {% endtabs %} {% columns %} {% column %} Nous pouvons maintenant connecter le serveur de modèles. Ouvrez \*\*Paramètres → Connexions\*\*, puis cliquez sur \*\*Ajouter un fournisseur\*\*. Sélectionnez llama.cpp, vLLM ou Ollama puis collez le \*\*URL de base\*\*. \* Exemple llama.cpp : \`http://localhost:8080/v1\` \* Exemple Ollama : \`http://localhost:11434/v1\` {% endcolumn %} {% column %} ![](https://unsloth.ai/files/585e33030c2495c534cf3e108171585135da48e4) {% endcolumn %} {% endcolumns %} Cliquez sur \*\*Charger les modèles\*\* pour récupérer les identifiants de modèle disponibles, ou saisissez-les manuellement si votre serveur n'expose pas \`/models\`. Puis, après avoir cliqué sur \*\*Ajouter un fournisseur,\*\* Les modèles que vous avez activés apparaîtront maintenant sous \*\*Externe\*\* dans le menu déroulant \*\*Sélectionner un modèle\*\* . ### Exécution de code Lorsqu'elle est activée, les modèles OpenAI et Anthropic pris en charge peuvent exécuter du code dans un bac à sable du fournisseur pour résoudre des problèmes, analyser des données et travailler avec des fichiers.\\ \\ Les modèles Anthropic utilisent l'outil d'exécution de code côté fournisseur de Claude. OpenAI utilise des conteneurs réutilisables, que vous pouvez créer, supprimer et sélectionner depuis les \*\*Exécution de code\*\* paramètres. Sélectionnez le même conteneur dans un nouveau fil pour continuer avec ses fichiers et son état. ![](https://unsloth.ai/files/b9189053500670039450008c1ac7e9b4a52ffcc9) \### Mise en cache des prompts La mise en cache des prompts réduit la latence et le coût lorsque les requêtes réutilisent le même long préfixe. Elle est prise en charge pour les fournisseurs et serveurs compatibles, notamment OpenAI, Anthropic et llama.cpp. Utilisez le \*\*paramètre de mise en cache des prompts\*\* dans le panneau latéral pour contrôler le comportement du cache pour les connexions prises en charge. ![](https://unsloth.ai/files/ba383fce34ff4f8414f62d082fe7d4a954f62a23) Pour llama.cpp, la mise en cache des prompts est activée par défaut et peut être désactivée au démarrage \`llama-server\` avec : \`\`\`bash --no-cache-prompt \`\`\` ### Recherche web et réflexion La recherche web côté fournisseur est disponible pour les modèles pris en charge d'OpenAI, Anthropic, OpenRouter, Mistral, Gemini et Kimi. Le contrôle Think s'adapte au modèle sélectionné : certains modèles utilisent un commutateur activé/désactivé, tandis que les modèles à effort de raisonnement utilisent des niveaux de réflexion spécifiques au modèle. ![](https://unsloth.ai/files/13d313bae2997c0089f52df2852153ce1e52e70e) \### Génération d'images Tout comme GPT et Gemini, Unsloth prend également en charge la génération d'images. Vous pouvez modifier directement une image en cliquant sur le bouton « Modifier l'image » et en saisissant une nouvelle invite pour l'affiner ou la régénérer. Les images sont générées automatiquement lorsqu'elles sont demandées, mais vous pouvez désactiver ce comportement. Un bouton de téléchargement est également disponible, vous permettant d'enregistrer l'image dans sa résolution d'origine maximale. ![](https://unsloth.ai/files/ac54d5c00a3c800fa87accbf18b78c9493ed96aa) ![](https://unsloth.ai/files/c13f9580d432538e1749676ffe05ec3542668ed4) \### Dépannage Si un fournisseur ne parvient pas à se connecter, vérifiez que la clé API appartient au fournisseur sélectionné et qu'elle donne accès au modèle que vous avez choisi. Si un modèle n'apparaît pas après avoir cliqué sur \*\*Recharger les modèles\*\*, il se peut qu'il ne soit pas disponible pour votre compte. Vous pouvez quand même utiliser la liste de modèles par défaut d'Unsloth ou choisir un autre modèle. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/integrations/connections.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/troubleshooting-and-faqs.md). # Dépannage et FAQ Si vous rencontrez toujours des problèmes avec les versions ou les dépendances, veuillez utiliser notre \[image Docker\](/docs/fr/commencer/install/docker.md) qui contiendra tout préinstallé. {% hint style="success" %} \*\*Essayez toujours de mettre à jour Unsloth si vous rencontrez des problèmes.\*\* \`pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth\_zoo\` {% endhint %} ### Affiner un nouveau modèle non pris en charge par Unsloth ? Unsloth fonctionne avec n’importe quel modèle pris en charge par \`transformers\`. Si un modèle ne figure pas dans nos téléchargements ou ne fonctionne pas immédiatement, il est généralement quand même pris en charge ; certains modèles plus récents peuvent simplement nécessiter un petit ajustement manuel en raison de nos optimisations. Dans la plupart des cas, vous pouvez activer la compatibilité en définissant \`trust\_remote\_code=True\` dans votre script de fine-tuning. Voici un exemple utilisant \[DeepSeek-OCR\](/docs/fr/modeles/tutorials/deepseek-ocr-how-to-run-and-fine-tune.md): from huggingface_hub import snapshot_download snapshot_download("unsloth/DeepSeek-OCR", local_dir = "deepseek_ocr") model, tokenizer = FastVisionModel.from_pretrained( "./deepseek_ocr", load_in_4bit = False, # Utiliser le 4bit pour réduire l’utilisation de la mémoire. False pour LoRA 16bit. auto_model = AutoModel, trust_remote_code = True, # Activer pour prendre en charge les nouveaux modèles unsloth_force_compile = True, use_gradient_checkpointing = "unsloth", # True ou "unsloth" pour long contexte ) \### L’exécution dans Unsloth fonctionne bien, mais après l’exportation et l’exécution sur d’autres plateformes, les résultats sont médiocres Il peut parfois arriver que votre modèle s’exécute et produise de bons résultats dans Unsloth, mais lorsque vous l’utilisez sur une autre plateforme comme Ollama ou vLLM, les résultats sont médiocres ou vous pouvez obtenir du charabia, des générations infinies/sans fin \*ou\* des sorties répétées\*\*.\*\* \* La cause la plus courante de cette erreur est l’utilisation d’un \*\*mauvais modèle de chat\*\*\*\*.\*\* Il est essentiel d’utiliser le MÊME modèle de chat qui a été utilisé lors de l’entraînement du modèle dans Unsloth et ensuite lorsque vous l’exécutez dans un autre framework, tel que llama.cpp ou Ollama. Lors de l’inférence à partir d’un modèle enregistré, il est crucial d’appliquer le bon modèle. \* Cela peut aussi être dû au fait que votre moteur d’inférence ajoute un jeton de « début de séquence » inutile (ou, à l’inverse, qu’il n’en ajoute pas), alors assurez-vous de vérifier les deux hypothèses ! \* \*\*Utilisez nos notebooks conversationnels pour forcer le modèle de chat - cela résoudra la plupart des problèmes.\*\* \* Notebook conversationnel Qwen-3 14B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_\\(14B\\)-Reasoning-Conversational.ipynb) \* Notebook conversationnel Gemma-3 4B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3\_\\(4B\\).ipynb) \* Notebook conversationnel Llama-3.2 3B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2\_\\(1B\_and\_3B\\)-Conversational.ipynb) \* Notebook conversationnel Phi-4 14B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi\_4-Conversational.ipynb) \* Notebook conversationnel Mistral v0.3 7B \[\*\*Ouvrir dans Colab\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral\_v0.3\_\\(7B\\)-Conversational.ipynb) \* \*\*Plus de notebooks dans notre\*\* \[\*\*documentation des notebooks\*\*\](/docs/fr/commencer/unsloth-notebooks.md) ### L’enregistrement en GGUF / vLLM 16 bits plante Vous pouvez essayer de réduire l’utilisation maximale du GPU pendant l’enregistrement en modifiant \`maximum\_memory\_usage\`. La valeur par défaut est \`model.save\_pretrained(..., maximum\_memory\_usage = 0.75)\`. Réduisez-la par exemple à 0,5 pour utiliser 50 % du pic de mémoire du GPU ou moins. Cela peut réduire les plantages OOM pendant l’enregistrement. ### Comment enregistrer manuellement en GGUF ? Enregistrez d’abord votre modèle en 16 bits via : \`\`\`python model.save\_pretrained\_merged("merged\_model", tokenizer, save\_method = "merged\_16bit",) \`\`\` Compilez llama.cpp à partir des sources comme ci-dessous : \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=ON -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split llama-mtmd-cli cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` Ensuite, enregistrez le modèle en F16 : \`\`\`bash python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-F16.gguf --outtype f16 \\ --split-max-size 50G \`\`\` \`\`\`bash # Pour BF16 : python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-BF16.gguf --outtype bf16 \\ --split-max-size 50G # Pour Q8\_0 : python llama.cpp/convert\_hf\_to\_gguf.py merged\_model \\ --outfile model-Q8\_0.gguf --outtype q8\_0 \\ --split-max-size 50G \`\`\` ### Pourquoi Q8\\\_K\\\_XL est-il plus lent que Q8\\\_0 GGUF ? Sur les appareils Mac, il semble que BF16 puisse être plus lent que F16. Q8\\\_K\\\_XL convertit certaines couches en BF16, d’où le ralentissement. Nous modifions activement notre processus de conversion afin de faire de F16 le choix par défaut pour Q8\\\_K\\\_XL et de réduire les pertes de performance. ### Comment faire l’évaluation Pour configurer l’évaluation dans votre exécution d’entraînement, vous devez d’abord répartir votre jeu de données en un ensemble d’entraînement et un ensemble de test. Vous devriez \*\*toujours mélanger la sélection du jeu de données\*\*, sinon votre évaluation est erronée ! \`\`\`python new\_dataset = dataset.train\_test\_split( test\_size = 0.01, # 1 % pour la taille de test peut aussi être un entier pour le # de lignes shuffle = True, # Devrait toujours être défini sur True ! seed = 3407, ) train\_dataset = new\_dataset\["train"\] # Jeu de données pour l’entraînement eval\_dataset = new\_dataset\["test"\] # Jeu de données pour l’évaluation \`\`\` Ensuite, nous pouvons définir les arguments d’entraînement pour activer l’évaluation. Rappel : l’évaluation peut être très, très lente, surtout si vous définissez \`eval\_steps = 1\` ce qui signifie que vous évaluez à chaque étape. Si c’est le cas, essayez de réduire la taille de eval\\\_dataset à, disons, 100 lignes ou quelque chose comme ça. \`\`\`python from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( args = SFTConfig( fp16\_full\_eval = True, # Définissez ceci pour réduire l’utilisation de la mémoire per\_device\_eval\_batch\_size = 2,# L’augmenter utilisera plus de mémoire eval\_accumulation\_steps = 4, # Vous pouvez augmenter cela à la place de la taille du lot eval\_strategy = "steps", # Lance l’évaluation toutes les quelques étapes ou époques. eval\_steps = 1, # Nombre d’évaluations effectuées par # d’étapes d’entraînement ), train\_dataset = new\_dataset\["train"\], eval\_dataset = new\_dataset\["test"\], ... ) trainer.train() \`\`\` ### Boucle d’évaluation - Manque de mémoire ou plantage. Un problème courant lorsque vous manquez de mémoire (OOM) est que vous avez défini une taille de lot trop élevée. Réglez-la à moins de 2 pour utiliser moins de VRAM. Utilisez aussi \`fp16\_full\_eval=True\` pour utiliser float16 pour l’évaluation, ce qui réduit la mémoire de moitié. Commencez par diviser votre jeu de données d’entraînement en un ensemble d’entraînement et un ensemble de test. Définissez les paramètres du formateur pour l’évaluation ainsi : \`\`\`python new\_dataset = dataset.train\_test\_split(test\_size = 0.01) from trl import SFTTrainer, SFTConfig trainer = SFTTrainer( args = SFTConfig( fp16\_full\_eval = True, per\_device\_eval\_batch\_size = 2, eval\_accumulation\_steps = 4, eval\_strategy = "steps", eval\_steps = 1, ), train\_dataset = new\_dataset\["train"\], eval\_dataset = new\_dataset\["test"\], ... ) \`\`\` Cela évitera les OOM et rendra le tout un peu plus rapide. Vous pouvez aussi utiliser \`bf16\_full\_eval=True\` pour les machines bf16. Par défaut, Unsloth devrait avoir défini ces drapeaux par défaut à partir de juin 2025. ### Comment puis-je faire de l'arrêt anticipé ? Si vous souhaitez arrêter l’exécution de fine-tuning / d’entraînement parce que la perte d’évaluation ne diminue pas, vous pouvez utiliser l’arrêt anticipé, qui arrête le processus d’entraînement. Utilisez \`EarlyStoppingCallback\`. Comme d'habitude, configurez votre trainer et votre ensemble de données de validation. Ce qui suit est utilisé pour arrêter l'exécution de l'entraînement si le \`eval\_loss\` (la perte de validation) ne diminue pas après environ 3 étapes. \`\`\`python from trl import SFTConfig, SFTTrainer trainer = SFTTrainer( args = SFTConfig( fp16\_full\_eval = True, per\_device\_eval\_batch\_size = 2, eval\_accumulation\_steps = 4, output\_dir = "training\_checkpoints", # emplacement des points de contrôle enregistrés pour l'arrêt anticipé save\_strategy = "steps", # enregistrer le modèle toutes les N étapes save\_steps = 10, # nombre d'étapes avant l'enregistrement du modèle save\_total\_limit = 3, # ne conserver que 3 points de contrôle enregistrés pour économiser de l'espace disque eval\_strategy = "steps", # évaluer toutes les N étapes eval\_steps = 10, # nombre d'étapes avant d'effectuer l'évaluation load\_best\_model\_at\_end = True, # DOIT ÊTRE UTILISÉ pour l'arrêt anticipé metric\_for\_best\_model = "eval\_loss", # métrique sur laquelle nous voulons appliquer l'arrêt anticipé greater\_is\_better = False, # plus la perte de validation est faible, mieux c'est ), model = model, tokenizer = tokenizer, train\_dataset = new\_dataset\["train"\], eval\_dataset = new\_dataset\["test"\], ) \`\`\` Nous ajoutons ensuite le callback, qui peut également être personnalisé : \`\`\`python from transformers import EarlyStoppingCallback early\_stopping\_callback = EarlyStoppingCallback( early\_stopping\_patience = 3, # Combien d'étapes nous attendrons si la perte de validation ne diminue pas # Par exemple, la perte peut augmenter, puis diminuer après 3 étapes early\_stopping\_threshold = 0.0, # Peut être défini plus haut - définit de combien la perte doit diminuer jusqu'à # ce que nous considérions un arrêt anticipé. Par exemple, 0.01 signifie que si la perte était # 0.02 puis 0.01, nous considérons qu'il faut arrêter l'exécution de manière anticipée. ) trainer.add\_callback(early\_stopping\_callback) \`\`\` Puis entraînez le modèle comme d'habitude via \`trainer.train() .\` ### Le téléchargement reste bloqué à 90 à 95 % Si votre modèle reste bloqué à 90, 95 % pendant longtemps, vous pouvez désactiver certains processus de téléchargement rapide afin de forcer des téléchargements synchrones et d’afficher davantage de messages d’erreur. Utilisez simplement \`UNSLOTH\_STABLE\_DOWNLOADS=1\` avant tout import d’Unsloth. \`\`\`python import os os.environ\["UNSLOTH\_STABLE\_DOWNLOADS"\] = "1" from unsloth import FastLanguageModel \`\`\` ### RuntimeError: erreur CUDA : assertion déclenchée côté périphérique Redémarrez et exécutez tout, mais placez ceci au début avant tout import d’Unsloth. Veuillez aussi déposer un rapport de bogue dès que possible, merci ! \`\`\`python import os os.environ\["UNSLOTH\_COMPILE\_DISABLE"\] = "1" os.environ\["UNSLOTH\_DISABLE\_FAST\_GENERATION"\] = "1" \`\`\` ### Toutes les étiquettes de votre jeu de données sont à -100. Les pertes d’entraînement seront toutes à 0. Cela signifie que votre utilisation de \`train\_on\_responses\_only\` est incorrecte pour ce modèle particulier. train\\\_on\\\_responses\\\_only vous permet de masquer la question de l’utilisateur et d’entraîner votre modèle à produire la réponse de l’assistant avec un poids plus élevé. Cela est connu pour augmenter la précision d’au moins 1 %. Consultez notre \[\*\*Guide des hyperparamètres LoRA\*\*\](/docs/fr/commencer/fine-tuning-llms-guide/lora-hyperparameters-guide.md) pour plus de détails. Pour les modèles de type Llama 3.1, 3.2, 3.3, veuillez utiliser ce qui suit : \`\`\`python from unsloth.chat\_templates import train\_on\_responses\_only trainer = train\_on\_responses\_only( trainer, instruction\_part = "<|start\_header\_id|>user<|end\_header\_id|>\\n\\n", response\_part = "<|start\_header\_id|>assistant<|end\_header\_id|>\\n\\n", ) \`\`\` Pour les modèles Gemma 2, 3, 3n, utilisez ce qui suit : \`\`\`python from unsloth.chat\_templates import train\_on\_responses\_only trainer = train\_on\_responses\_only( trainer, instruction\_part = "user\\n", response\_part = "model\\n", ) \`\`\` ### Unsloth est plus lent que prévu ? Si votre vitesse semble plus lente au début, c’est probablement parce que \`torch.compile\` met généralement environ 5 minutes (ou plus) à se réchauffer et à terminer la compilation. Assurez-vous de mesurer le débit \*\*après\*\* qu’il est complètement chargé, car sur des exécutions plus longues, Unsloth devrait être beaucoup plus rapide. Pour désactiver, utilisez : \`\`\`python import os os.environ\["UNSLOTH\_COMPILE\_DISABLE"\] = "1" \`\`\` ### Certaines poids de Gemma3nForConditionalGeneration n’ont pas été initialisés à partir du point de contrôle du modèle C’est une erreur critique, car cela signifie que certains poids ne sont pas analysés correctement, ce qui entraînera des sorties incorrectes. Cela peut normalement être corrigé en mettant à jour Unsloth \`pip install --upgrade --force-reinstall --no-cache-dir --no-deps unsloth unsloth\_zoo\` Puis mettez à jour transformers et timm : \`pip install --upgrade --force-reinstall --no-cache-dir --no-deps transformers timm\` Cependant, si le problème persiste, veuillez déposer un rapport de bogue dès que possible ! ### NotImplementedError : un environnement linguistique UTF-8 est requis. ANSI obtenu Voir Dans une nouvelle cellule, exécutez ce qui suit : \`\`\`python import locale locale.getpreferredencoding = lambda: "UTF-8" \`\`\` ### Citer Unsloth Si vous citez l’utilisation de nos téléchargements de modèles, utilisez le Bibtex ci-dessous. Ceci concerne Qwen3-30B-A3B-GGUF Q8\\\_K\\\_XL : \`\`\` @misc{unsloth\_2025\_qwen3\_30b\_a3b, author = {Unsloth AI and Han-Chen, Daniel and Han-Chen, Michael}, title = {Qwen3-30B-A3B-GGUF:Q8\\\_K\\\_XL}, year = {2025}, publisher = {Hugging Face}, howpublished = {\\url{https://huggingface.co/unsloth/Qwen3-30B-A3B-GGUF}} } \`\`\` Pour citer l’utilisation de notre package Github ou notre travail en général : \`\`\` @misc{unsloth, author = {Unsloth AI and Han-Chen, Daniel and Han-Chen, Michael}, title = {Unsloth}, year = {2025}, publisher = {Github}, howpublished = {\\url{https://github.com/unslothai/unsloth}} } \`\`\` \[^1\]: Activez cette ligne de code et voyez si cela fonctionne. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/troubleshooting-and-faqs.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/qwen3.5/fine-tune.md). # Guide de fine-tuning Qwen3.5 Vous pouvez maintenant effectuer un fine-tuning \[Qwen3.5\](/docs/fr/modeles/qwen3.5.md) la famille de modèles (0,8B, 2B, 4B, 9B, 27B, 35B‑A3B, 122B‑A10B) avec \[\*\*Unsloth\*\*\](https://github.com/unslothai/unsloth). La prise en charge inclut à la fois \[la vision\](/docs/fr/modeles/qwen3.5/fine-tune.md#vision-fine-tuning), le texte et \[RL\](#reinforcement-learning-rl) le fine-tuning. \*\*Qwen3.5‑35B‑A3B\*\* - le LoRA bf16 fonctionne sur \*\*74 Go de VRAM.\*\* \* Unsloth permet à Qwen3.5 de s’entraîner \*\*1,5× plus vite\*\* et utilise \*\*50 % de VRAM en moins\*\* que les configurations FA2. \* Utilisation de la VRAM pour le LoRA bf16 de Qwen3.5 : \*\*0,8B\*\*: 3 Go • \*\*2B\*\*: 5 Go • \*\*4B\*\*: 10 Go • \*\*9B\*\*: 22 Go • \*\*27B\*\*: 56 Go \* Effectuez un fine-tuning \*\*0,8B\*\*, \*\*2B\*\* et \*\*4B\*\* bf16 LoRA via nos \*\*notebooks Google Colab\*\* \*\*gratuits\*\*: | \[Qwen3.5-\*\*0,8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) | \[Qwen3.5-4B \*\*GRPO\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision\_GRPO.ipynb) | | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | \* Si vous voulez \*\*préserver la capacité de raisonnement\*\* vous pouvez mélanger des exemples de style raisonnement avec des réponses directes (conservez au minimum 75 % de raisonnement). Sinon, vous pouvez l’émettre entièrement. \* \*\*Le fine-tuning complet (FFT)\*\* fonctionne aussi. Notez qu’il utilisera 4× plus de VRAM. \* Qwen3.5 est puissant pour le fine-tuning multilingue, car il prend en charge 201 langues. \* Après le fine-tuning, vous pouvez exporter vers \[GGUF\](#saving-export-your-fine-tuned-model) (pour llama.cpp/Ollama/etc.) ou \[vLLM\](#saving-export-your-fine-tuned-model) \* \[Apprentissage par renforcement\](/docs/fr/commencer/reinforcement-learning-rl-guide.md) (RL) pour Qwen3.5 \[Le RL VLM\](/docs/fr/commencer/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md) fonctionne aussi via l’inférence Unsloth. \* Nous avons des notebooks \*\*A100\*\* Colab pour \[Qwen3.5‑27B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen\_3\_5\_27B\_A100\\(80GB\\).ipynb) et \[Qwen3.5‑35B‑A3B\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_MoE.ipynb). Si vous utilisez une version plus ancienne (ou si vous effectuez un fine-tuning en local), mettez d’abord à jour : {% columns %} {% column width="50%" %} Unsloth Studio : {% code expandable="true" %} \`\`\`bash unsloth studio update \`\`\` {% endcode %} {% endcolumn %} {% column width="50%" %} Unsloth basé sur le code : \`\`\`bash pip install --upgrade --force-reinstall --no-cache-dir unsloth unsloth\_zoo \`\`\` {% endcolumn %} {% endcolumns %} {% hint style="warning" %} \*\*Veuillez utiliser \`transformers v5\` pour Qwen3.5. Les versions plus anciennes ne fonctionneront pas. Unsloth utilise désormais automatiquement transformers v5 par défaut (sauf pour les environnements Colab).\*\* Si l’entraînement semble \*\*plus lent que d’habitude\*\*, c’est parce que Qwen3.5 utilise des noyaux Triton Mamba personnalisés. La compilation de ces noyaux peut prendre plus de temps que la normale, en particulier sur les GPU T4. Il n’est pas recommandé d’effectuer un entraînement QLoRA (4 bits) sur les modèles Qwen3.5, qu’ils soient MoE ou denses, en raison de différences de quantification plus élevées que la normale. {% endhint %} ### Fine-tuning MoE (35B, 122B) Pour les modèles MoE comme \*\*Qwen3.5‑35B‑A3B / 122B‑A10B / 397B‑A17B\*\*: \* Vous pouvez utiliser notre \[Qwen3.5‑35B‑A3B (A100)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_MoE.ipynb) notebook de fine-tuning \* Prend en charge notre récente mise à jour d’entraînement MoE \[\\~12× plus rapide\](/docs/fr/notions-de-base/faster-moe.md) avec >35 % de VRAM en moins et \\~6× plus de longueur de contexte \* \*\*Il est préférable d’utiliser des configurations bf16 (par ex. LoRA ou fine-tuning complet)\*\* (le MoE QLoRA 4 bits n’est pas recommandé en raison des limitations de BitsandBytes). \* Les noyaux MoE d’Unsloth sont activés par défaut et peuvent utiliser différents backends ; vous pouvez changer avec \`UNSLOTH\_MOE\_BACKEND\`. \* Le fine-tuning de la couche routeur est désactivé par défaut pour des raisons de stabilité. \* Qwen3.5‑122B‑A10B - le LoRA bf16 fonctionne sur 256 Go de VRAM. Si vous utilisez plusieurs GPU, ajoutez \`device\_map = "balanced"\` ou suivez notre \[Guide multiGPU\](/docs/fr/notions-de-base/multi-gpu-training-with-unsloth.md). ### Démarrage rapide #### 🦥 Guide d’Unsloth Studio Qwen3.5 peut être exécuté et affiné dans \[Unsloth Studio\](/docs/fr/nouveau/studio.md), notre nouvelle interface web open source pour l’IA locale. Avec Unsloth Studio, vous pouvez exécuter des modèles localement sur \*\*MacOS, Windows\*\*, Linux et : {% columns %} {% column %} \* \[Entraîner des LLMs\](/docs/fr/nouveau/studio.md#no-code-training) 2x plus rapide avec 70 % de VRAM en moins \* Rechercher, télécharger, \[exécuter des GGUFs\](/docs/fr/nouveau/studio.md#run-models-locally) et des modèles safetensor \* \[\*\*Auto-réparation\*\* appel d’outils\](/docs/fr/nouveau/studio.md#execute-code--heal-tool-calling) + \*\*recherche web\*\* \* \[\*\*Exécution de code\*\*\](/docs/fr/nouveau/studio.md#run-models-locally) (Python, Bash) \* \[Inférence automatique\](/docs/fr/nouveau/studio.md#model-arena) réglage des paramètres (temp, top-p, etc.) \* Inférence rapide CPU + GPU via llama.cpp {% endcolumn %} {% column %} ![](https://unsloth.ai/files/c8afb6876b3c9584c47f3bf91590355153005596) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} #### Installer Unsloth Exécutez dans votre terminal : \*\*MacOS, Linux, WSL :\*\* \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` \*\*Windows PowerShell :\*\* \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` {% hint style="success" %} \*\*L’installation sera rapide et prendra environ 1 à 2 minutes.\*\* {% endhint %} {% endstep %} {% step %} #### Lancer Unsloth \*\*MacOS, Linux, WSL et Windows :\*\* \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` \*\*Puis ouvrez \`http://localhost:8888\` dans votre navigateur.\*\* {% endstep %} {% step %} #### Entraîner Qwen3.5 Au premier lancement, vous devrez créer un mot de passe pour sécuriser votre compte et vous reconnecter plus tard. Vous verrez ensuite un bref assistant d’intégration pour choisir un modèle, un jeu de données et des paramètres de base. Vous pouvez le passer à tout moment. Recherchez Qwen3.5 dans la barre de recherche et sélectionnez le modèle et le jeu de données souhaités. Ensuite, ajustez vos hyperparamètres et la longueur de contexte selon vos besoins. ![](https://unsloth.ai/files/121e202bfe70092d671bd4bd2c4fd4b52c503808) {% endstep %} {% step %} #### Surveiller la progression de l’entraînement Après avoir cliqué sur démarrer l’entraînement, vous pourrez surveiller et observer la progression de l’entraînement du modèle. La perte d’entraînement devrait diminuer régulièrement.\\ Une fois terminé, le modèle sera automatiquement enregistré. ![](https://unsloth.ai/files/522c87fa7f68badff0a2575faa7f20b7aeb919ac) {% endstep %} {% step %} #### Exporter votre modèle affiné Une fois terminé, Unsloth Studio vous permet d’exporter le modèle vers les formats GGUF, safetensor, etc. ![](https://unsloth.ai/files/e7aa65e61a9167698dca0232b2f5d3b23952c509) {% endstep %} {% endstepper %} #### Guide d’Unsloth Core (basé sur le code) : Voici une recette SFT minimale (fonctionne pour le fine-tuning « texte uniquement »). Voir aussi notre \[fine-tuning de la vision\](/docs/fr/notions-de-base/vision-fine-tuning.md) section. {% hint style="info" %} Qwen3.5 est un « modèle de langage causal avec encodeur de vision » (c’est un VLM unifié), donc assurez-vous d’avoir les dépendances vision habituelles installées (\`torchvision\`, \`pillow\`) si nécessaire, et maintenez Transformers à jour. Utilisez la dernière version de Transformers pour Qwen3.5. \*\*Si vous souhaitez faire\*\* \[\*\*GRPO\*\*\](/docs/fr/commencer/reinforcement-learning-rl-guide.md)\*\*, cela fonctionne dans Unsloth si vous désactivez l’inférence rapide vLLM et utilisez à la place l’inférence Unsloth. Suivez nos exemples de notebook\*\* \[\*\*Vision RL\*\*\](/docs/fr/commencer/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl.md) \*\*.\*\* {% endhint %} {% code expandable="true" %} \`\`\`python from unsloth import FastLanguageModel import torch from datasets import load\_dataset from trl import SFTTrainer, SFTConfig max\_seq\_length = 2048 # commencez petit ; augmentez après que cela fonctionne # Jeu de données d’exemple (remplacez-le par le vôtre). Nécessite une colonne "text". url = "https://huggingface.co/datasets/laion/OIG/resolve/main/unified\_chip2.jsonl" dataset = load\_dataset("json", data\_files={"train": url}, split="train") model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "Qwen/Qwen3.5-27B", max\_seq\_length = max\_seq\_length, load\_in\_4bit = False, # Le QLoRA MoE n’est pas recommandé, le dense 27B fonctionne bien load\_in\_16bit = True, # LoRA bf16/16 bits full\_finetuning = False, ) model = FastLanguageModel.get\_peft\_model( model, r = 16, target\_modules = \[ "q\_proj", "k\_proj", "v\_proj", "o\_proj", "gate\_proj", "up\_proj", "down\_proj", \], lora\_alpha = 16, lora\_dropout = 0, bias = "none", # Le checkpointing "unsloth" est conçu pour un très long contexte + une VRAM plus faible use\_gradient\_checkpointing = "unsloth", random\_state = 3407, max\_seq\_length = max\_seq\_length, ) trainer = SFTTrainer( model = model, train\_dataset = dataset, tokenizer = tokenizer, args = SFTConfig( max\_seq\_length = max\_seq\_length, per\_device\_train\_batch\_size = 1, gradient\_accumulation\_steps = 4, warmup\_steps = 10, max\_steps = 100, logging\_steps = 1, output\_dir = "outputs\_qwen35", optim = "adamw\_8bit", seed = 3407, dataset\_num\_proc = 1, ), ) trainer.train() \`\`\` {% endcode %} {% hint style="info" %} Si vous avez une erreur OOM : \* Diminuez \`per\_device\_train\_batch\_size\` par \*\*1\*\* et/ou réduisez \`max\_seq\_length\`. \* Conservez \`use\_\`\[\`gradient\_checkpointing\`\](/docs/fr/blog/500k-context-length-fine-tuning.md#unsloth-gradient-checkpointing-enhancements)\`="unsloth"\` activé (il est conçu pour réduire l’utilisation de la VRAM et étendre la longueur de contexte). {% endhint %} \*\*Exemple de chargeur pour MoE (LoRA bf16) :\*\* \`\`\`python import os import torch from unsloth import FastModel model, tokenizer = FastModel.from\_pretrained( model\_name = "unsloth/Qwen3.5-35B-A3B", max\_seq\_length = 2048, load\_in\_4bit = False, # Le QLoRA MoE n’est pas recommandé, le dense 27B fonctionne bien load\_in\_16bit = True, # LoRA bf16/16 bits full\_finetuning = False, ) \`\`\` Une fois chargé, vous attacherez des adaptateurs LoRA et entraînerez de manière similaire à l’exemple SFT ci-dessus. ### Fine-tuning de la vision Unsloth prend en charge \[fine-tuning de la vision\](/docs/fr/notions-de-base/vision-fine-tuning.md) pour les modèles multimodaux Qwen3.5. Utilisez les notebooks Qwen3.5 ci-dessous et remplacez les noms de modèles respectifs par le modèle Qwen3.5 souhaité. | \[Qwen3.5-\*\*0,8B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(0\_8B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*2B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(2B\\)\_Vision.ipynb) | \[Qwen3.5-\*\*4B\*\*\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision.ipynb) | Qwen3.5-\*\*9B\*\* | | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | -------------- | \* \[Notebook RL GRPO/GSPO Qwen3-VL\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_VL\_\\(8B\\)-Vision-GRPO.ipynb) (changez le nom du modèle en Qwen3.5-4B, etc.) \*\*Désactivation du fine-tuning Vision / texte uniquement :\*\* Pour effectuer un fine-tuning des modèles de vision, nous vous permettons désormais de sélectionner quelles parties du modèle affiner. Vous pouvez choisir de n’affiner que les couches de vision, ou les couches de langage, ou les couches d’attention / MLP ! Nous les activons toutes par défaut ! {% code expandable="true" %} \`\`\`python model = FastVisionModel.get\_peft\_model( model, finetune\_vision\_layers = True, # False si vous n’affinez pas les couches de vision finetune\_language\_layers = True, # False si vous n’affinez pas les couches de langage finetune\_attention\_modules = True, # False si vous n’affinez pas les couches d’attention finetune\_mlp\_modules = True, # False si vous n’affinez pas les couches MLP r = 16, # Plus la valeur est grande, plus la précision est élevée, mais cela peut surajuster lora\_alpha = 16, # Alpha recommandé = r au minimum lora\_dropout = 0, bias = "none", random\_state = 3407, use\_rslora = False, # Nous prenons en charge le LoRA à rang stabilisé loftq\_config = None, # Et LoftQ target\_modules = "all-linear", # Optionnel désormais ! Vous pouvez spécifier une liste si nécessaire modules\_to\_save=\[ "lm\_head", "embed\_tokens", \], ) \`\`\` {% endcode %} Afin d’affiner ou d’entraîner Qwen3.5 avec plusieurs images, consultez notre \[\*\*guide de vision multi-images\*\*\](/docs/fr/notions-de-base/vision-fine-tuning.md#multi-image-training)\*\*.\*\* ### Apprentissage par renforcement (RL) Vous pouvez maintenant entraîner Qwen3.5 avec RL, GSPO, GRPO, etc. avec \[notre notebook gratuit\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3\_5\_\\(4B\\)\_Vision\_GRPO.ipynb): {% embed url="" %} Vous pouvez exécuter le RL de Qwen3.5 avec Unsloth même s’il n’est pas pris en charge par vLLM, en définissant \`fast\_inference=False\` lors du chargement du modèle : \`\`\`python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from\_pretrained( model\_name="unsloth/Qwen3.5-4B", fast\_inference=False, ) \`\`\` ### Enregistrement / exportation du modèle affiné Vous pouvez consulter nos guides spécifiques d’inférence / déploiement pour \[Unsloth Studio\](/docs/fr/nouveau/studio/export.md), \[llama.cpp\](/docs/fr/notions-de-base/inference-and-deployment/saving-to-gguf.md), \[vLLM\](/docs/fr/notions-de-base/inference-and-deployment/vllm-guide.md), \[llama-server\](/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint.md), \[Ollama\](/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama.md). #### Enregistrer en GGUF Unsloth prend en charge l’enregistrement direct en GGUF : \`\`\`python model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q4\_k\_m") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "q8\_0") model.save\_pretrained\_gguf("directory", tokenizer, quantization\_method = "f16") \`\`\` Ou pousser les GGUF vers Hugging Face : \`\`\`python model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q4\_k\_m") model.push\_to\_hub\_gguf("hf\_username/directory", tokenizer, quantization\_method = "q8\_0") \`\`\` Si le modèle exporté se comporte moins bien dans un autre environnement d’exécution, Unsloth identifie la cause la plus courante : \*\*mauvais modèle de chat / jeton EOS au moment de l’inférence\*\* (vous devez utiliser le même modèle de chat que celui utilisé pour l’entraînement). #### Enregistrer pour vLLM {% hint style="warning" %} version vLLM \`0.16.0\` ne prend pas en charge Qwen3.5. Attendez jusqu’à \`0.170\` ou essayez la version Nightly. {% endhint %} Pour enregistrer en 16 bits pour vLLM, utilisez : {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("finetuned\_model", tokenizer, save\_method = "merged\_16bit") ## OU pour téléverser sur Hugging Face : model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "merged\_16bit", token = "") \`\`\` {% endcode %} Pour enregistrer uniquement les adaptateurs LoRA, utilisez soit : \`\`\`python model.save\_pretrained("finetuned\_lora") tokenizer.save\_pretrained("finetuned\_lora") \`\`\` Ou utilisez notre fonction intégrée : {% code overflow="wrap" %} \`\`\`python model.save\_pretrained\_merged("finetuned\_model", tokenizer, save\_method = "lora") ## OU pour téléverser sur Hugging Face model.push\_to\_hub\_merged("hf/model", tokenizer, save\_method = "lora", token = "") \`\`\` {% endcode %} Pour plus de détails, lisez nos guides d’inférence : {% columns %} {% column width="50%" %} {% content-ref url="/pages/44b6f06033c7dbf3b6521a33337058e295acc604" %} \[Inférence et déploiement\](/docs/fr/notions-de-base/inference-and-deployment.md) {% endcontent-ref %} {% content-ref url="/pages/0ce33fc68eed069d43cdcfb76b9793ce71c64c1f" %} \[GGUF & llama.cpp\](/docs/fr/notions-de-base/inference-and-deployment/saving-to-gguf.md) {% endcontent-ref %} {% endcolumn %} {% column width="50%" %} {% content-ref url="/pages/817a1275219e1e8d86fe100d223ac4b0862ab3a1" %} \[Model Export\](/docs/fr/nouveau/studio/export.md) {% endcontent-ref %} {% content-ref url="/pages/682151c53afcf1f6d611eb29ad62b7182b5187ea" %} \[vLLM\](/docs/fr/notions-de-base/inference-and-deployment/vllm-guide.md) {% endcontent-ref %} {% endcolumn %} {% endcolumns %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/qwen3.5/fine-tune.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/nvfp4.md). # Guide d'exécution d'Unsloth Dynamic NVFP4 Unsloth Dynamic NVFP4 est un format de modèle quantifié qui s’exécute sur les GPU NVIDIA Blackwell et est conçu pour une inférence 4 bits plus rapide et plus précise. Il combine la précision NVFP4 native de NVIDIA avec \[Unsloth Dynamic 2.0\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) la quantification pour préserver la précision du modèle tout en réduisant l’utilisation de la VRAM et en augmentant la vitesse. Ce guide explique la quantification FP4, compare NVFP4 avec d’autres formats et montre comment exécuter des modèles comme \[Gemma 4\](/docs/fr/modeles/gemma-4.md) et \[Qwen3.6\](/docs/fr/modeles/qwen3.6.md) localement à l’aide de vLLM ou SGLang sur les RTX 5050-5090, B200, RTX PRO 6000 et d’autres GPU. Dynamic NVFP4 fonctionne en sélectionnant les couches importantes pour rester en FP8 (W8A8) ou BF16 et le reste en W4A4 (et non W4A16) au lieu de forcer chaque couche en FP4. Cela permet jusqu’à \*\*une inférence 2,5x plus rapide\*\* car W4A4 exploite les cœurs tensoriels FP4 du GPU Blackwell. Pour toutes les quantifications, nous fournissons également un calibrage du cache KV en FP8 permettant \*\*des longueurs de contexte 2x plus longues\*\*. {% hint style="success" %} \*\*Tous\*\* \[\*\*Gemma 4\*\*\](#gemma-4) \*\*les modèles sont désormais disponibles sous forme de quantifications Unsloth Dynamic NVFP4 :\*\* E2B, E4B, 12B Unified, 26B-A4B MoE et 31B Dense. Découvrez la \[collection Unsloth Dynamic NVFP4\](https://huggingface.co/collections/unsloth/nvfp4) pour tous nos téléversements de modèles. {% endhint %} ### Float4 vs autres précisions L’astuce pour \*\*les GPU plus rapides consiste à réduire la précision numérique des multiplications matricielles\*\*. Le nombre de transistors nécessaires pour les unités de multiplication matricielle est lié au \*\*carré de la mantisse\*\*. La mantisse permet aux nombres d’avoir combien de décimales « fractionnaires » - donc plus il y a de bits, plus il peut représenter les décimales avec précision. Par exemple, exprimer 0.121332 est possible avec plus de bits de mantisse, tandis que peu de bits de mantisse le ramèneront à 0.1. {% columns %} {% column width="50%" %} FP32 a 23 bits de mantisse, donc 23^2 + 8 bits d’exposant = il faut 537 unités d’espace. Le bfloat16 a 7 bits de mantisse, donc 7^2 + 8 bits d’exposant = il faut 57 unités d’espace. Cela signifie que le bfloat16 nécessite environ 9x moins d’espace que le FP32 ! Et quand on passe au float8, qui a 3 bits de mantisse, donc 3^2 + 4 bits d’exposant = 13 - c’est 41x moins d’espace que le FP32 ! Enfin, float4 a 1 bit de mantisse et 2 exposants donc 3 unités d’espace - un énorme 179x moins d’espace que FP32 - cela signifie essentiellement qu’un \*\*GPU peut faire environ 179x plus de multiplications matricielles FP4 que de FLOPs de multiplication FP32 dans le même espace\*\*! {% endcolumn %} {% column width="50%" %} !\[\](/files/9d3b2e37418758847bd366f02b17b30475ad8bc2) {% endcolumn %} {% endcolumns %} ### NVFP4 vs MXFP4 ![](https://unsloth.ai/files/0895729245aa3e7c0e78200d3c6adad067aafdfd) ![](https://unsloth.ai/files/455b2e156c993ae94ad8449f52c28e803856fc9d) Il existe un autre format FP4 appelé MXFP4 - il est moins précis que NVFP4 pour 2 raisons : 1. NVFP4 utilise une taille de bloc de 16 contre 32 pour MXFP4 - cela permet d’isoler plus facilement les valeurs aberrantes et des facteurs d’échelle sont fournis pour des sous-ensembles plus petits de poids, ce qui augmente la précision 2. Une échelle E4M3 (FP8) est utilisée au lieu d’une E8M0 (mise à l’échelle par puissances de 2) par bloc. L’utilisation d’une taille de bloc de type FP8 semble bien meilleure, surtout pour les LLM. ### Analyse des performances Nos nouvelles quantifications dynamiques NVFP4 de Qwen3.6 s’exécutent \\~\*\*2,5x plus vite\*\* que les autres quantifications NVFP4, avec \*\*de meilleures performances\*\* et des tailles de fichier comparables. Exécutez Qwen3.6-27B NVFP4 \*\*2,5x plus vite\*\* sur \*\*24 Go de VRAM\*\* et Qwen3.6-35B-A3B \*\*1,7x plus vite\*\* sur \*\*32 Go de VRAM\*\*. Nous avons également ajouté \*\*le calibrage du cache KV en FP8\*\* pour des longueurs de contexte 2x plus longues ! NVFP4 nécessite les GPU Blackwell de NVIDIA comme les RTX 50X, DGX Spark (voir \[#dgx-spark-with-nvfp4-quants\](#dgx-spark-with-nvfp4-quants "mention")), B200, B300 GPU. Pour les GPU plus anciens, nos GGUF fonctionnent bien ! ![](https://unsloth.ai/files/6cab20b335b1ed16dbd30891fdeab4340f29737d) Tous les benchmarks utilisent 1x B200 en 128 de concurrence. Une concurrence plus élevée peut porter le 35B à 17 561 tokens / s. Nous publions également deux versions NVFP4 35B-A3B : \* \[Qwen3.6-35B-A3B-NVFP4-Fast\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) qui est une quantification W4A4 complète - 1,79x plus rapide \* \[Qwen3.6-35B-A3B-NVFP4\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4) qui est légèrement plus volumineuse mais plus précise et 1,56x plus rapide Pour les benchmarks de précision, nous avons mené MMLU-Pro, AIME 2025, GPQA pour FP8, BF16, le NVFP4 de NVIDIA et nos NVFP4 - nous montrons que nos quantifications plus rapides se comportent de manière similaire sur tout : ![](https://unsloth.ai/files/9621db18863d348297c1afba456da75abb3a4560) | Qwen3.6-35B-A3B | Qwen3.6-27B | | --- | --- | | [Qwen3.6-35B-A3B-NVFP4](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4)
(1,56x plus rapide) | [Qwen3.6-27B-NVFP4](https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4)
(2,5x plus rapide) | | [Qwen3.6-35B-A3B-NVFP4-Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast)
(1,79x plus rapide) | | \*\*Les tenseurs MTP sont également intégrés directement dans les quantifications pour des gains de vitesse supplémentaires.\*\* Les gains de précision proviennent d’améliorations apportées au modèle de conversation et au calibrage du jeu de données de Qwen3.6. Nous utilisons nos précédentes mises à jour du modèle de conversation pour améliorer la cohérence du codage et des appels d’outils tout en réduisant les boucles et d’autres problèmes signalés. Notre calibrage utilise un mélange de notre jeu de données optimisé pour le codage, les appels d’outils et le chat, ainsi qu’UltraChat. Pour la vitesse de décodage (tokens par personne), le nôtre est 1,03x plus rapide pour 27B et 1,17x et 1,22x plus rapide pour 35B. ![](https://unsloth.ai/files/f9e1fe845c177fc9c7517a05b98a3d7540b8909f) \### Vue d’ensemble Ci-dessous se trouvent les exigences matérielles pour les modèles que vous pouvez utiliser, notamment Gemma 4 et Qwen3.6. Voyez aussi le gain de vitesse global que vous obtiendrez : #### Gemma 4 : | variante de Gemma 4 | VRAM requise | Plus rapide que BF16 | | ------------------------------------------------------------------ | -----------: | -------------------: | | \[E2B\](https://huggingface.co/unsloth/gemma-4-E2B-it-NVFP4) | 7 Go | 1,12× plus rapide | | \[E4B\](https://huggingface.co/unsloth/gemma-4-E4B-it-NVFP4) | 9 Go | 1,22× plus rapide | | \[12B Unified\](https://huggingface.co/unsloth/gemma-4-12b-it-NVFP4) | 11 Go | 1,26× plus rapide | | \[26B A4B\](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-NVFP4) | 26 Go | 1,41× plus rapide | | \[31B\](https://huggingface.co/unsloth/gemma-4-31B-it-NVFP4) | 32 Go | 1,45× plus rapide | ![](https://unsloth.ai/files/8c1bf8707b4d15b5d6b80b93854724dad339e0cc) \#### Qwen3.6 : | variante de Qwen3.6 | VRAM requise | Plus rapide que les autres quantifications NVFP4 | | ------------------------------------------------------------------------- | -----------: | -----------------------------------------------: | | \[27B\](https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4) | 24 Go | 2,5x plus vite | | \[35B A3B\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4) | 32 Go | 1,56× plus rapide | | \[35B A3B Fast\](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) | 32 Go | 1,79× plus rapide | ### Benchmarks NVFP4 NVFP4 exécute directement les poids 4 bits et les multiplications matricielles sur les cœurs tensoriels Blackwell. Nos quantifications NVFP4 de Qwen3.6 utilisent W4A4, donc elles utilisent réellement les cœurs tensoriels FP4, et décodent donc plus vite que celles de NVIDIA qui utilisent W4A16. Nous quantifions également dynamiquement les couches pour conserver la précision, et nous avons réalisé MMLU-Pro, AIME 2025, GPQA pour toutes les quantifications, y compris la comparaison avec FP8 et BF16. \*\*Benchmarks de précision Qwen3.6-27B NVFP4\*\* | Fournisseur | MMLU-Pro | GPQA | AIME 2025 | | ----------- | -------: | ----: | --------: | | Unsloth | 86.25 | 86.34 | 93.12 | | NVIDIA | 85.96 | 86.87 | 93.12 | | FP8 | 86.11 | 86.87 | 93.75 | | BF16 | 85.96 | 88.13 | 93.33 | \*\*Benchmarks de précision Qwen3.6-35B-A3B NVFP4\*\* | Fournisseur | MMLU-Pro | GPQA | AIME 2025 | | ---------------- | -------: | ----: | --------: | | Unsloth | 85.85 | 86.74 | 92.29 | | \*\*Unsloth Fast\*\* | 85.58 | 87.75 | 91.67 | | NVIDIA | 85.60 | 87.12 | 91.88 | | FP8 | 85.75 | 86.74 | 93.12 | | BF16 | 85.75 | 86.36 | 92.50 | Nous avons également vérifié la longueur de sortie de tous les benchmarks, et elles sont comparables, donc les nouvelles quantifications NVFP4 ne réfléchissent pas plus longtemps, ce qui annulerait le but de les quantifier ! (C.-à-d. si c’est 2x plus rapide, mais que ça réfléchit 2x plus, alors c’est inutile) ![](https://unsloth.ai/files/35c7d92d632356b43389fe5d0180165634f89ebd) \## \*\*Exécuter les tutoriels NVFP4\*\* Pour exécuter les quantifications NVFP4, voir ci-dessous les commandes pour exécuter Qwen3.6-27B dans \[vLLM\](/docs/fr/notions-de-base/inference-and-deployment/vllm-guide.md) et \[SGLang\](/docs/fr/notions-de-base/inference-and-deployment/sglang-guide.md) (vous pouvez changer le nom du modèle pour \`Qwen3.6-35-A3B-NVFP4\`). ### \*\*Tutoriel vLLM\*\* Vous pouvez exécuter tous les modèles NVFP4 dans \[vLLM\](https://github.com/vllm-project/vllm). NE sélectionnez PAS de backend MoE - laissez vLLM le choisir - par exemple Marlin est 2,5x plus lent ! Voir \[#marlin-vs-flashinfer-vs-cutlass-vs-cute-dsl\](#marlin-vs-flashinfer-vs-cutlass-vs-cute-dsl "mention")Si vous avez un DGX Spark, voir \[#dgx-spark-serving\](#dgx-spark-serving "mention") vous devez utiliser \`--moe-backend flashinfer\_b12x\` sinon, l’inférence sera beaucoup plus lente. Pour installer vLLM dans un venv séparé : {% code overflow="wrap" expandable="true" %} \`\`\`bash uv venv unsloth-nvfp4-env --python 3.13 source unsloth-nvfp4-env/bin/activate uv pip install "vllm>=0.25.0" "flashinfer-python>=0.6.13" "nvidia-cutlass-dsl>=4.5.2" \\ --torch-backend=auto \`\`\` {% endcode %} Puis pour servir la variante 35B Fast : \`\`\`shell vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast \`\`\` Remplacez \`unsloth/Qwen3.6-35B-A3B-NVFP4-Fast\` par les noms de quantification NVFP4 ! Pour activer MTP / le décodage spéculatif (décodage plus rapide mais débit un peu plus faible), utilisez : \`\`\`bash vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast --speculative-config '{"method": "mtp", "num\_speculative\_tokens": 2}' \`\`\` Si vous rencontrez des problèmes avec Torchcodec, assurez-vous d’effectuer l’étape ci-dessous puis de relancer vllm. {% code overflow="wrap" expandable="true" %} \`\`\`bash sudo apt-get update sudo apt-get install -y ffmpeg \`\`\` {% endcode %} ### \*\*Tutoriel DGX Spark\*\* Pour garantir que le DGX Spark dispose des bons noyaux (sinon vous obtiendrez \*\*une inférence 2x PLUS LENTE\*\*), vérifiez d’abord : {% code overflow="wrap" expandable="true" %} \`\`\`bash python -c " import torch; from vllm.utils.flashinfer import has\_flashinfer\_b12x\_gemm as g, has\_flashinfer\_b12x\_moe as m cap = torch.cuda.get\_device\_capability(); print('cap', cap, '| b12x gemm', g(), '| b12x moe', m()); assert cap\[0\] == 12 and g() and m(), 'b12x indisponible : le service retomberait sur marlin W4A16'" \`\`\` {% endcode %} qui ne devrait PAS produire d’erreur - si c’est le cas, veuillez mettre à jour vllm ou réinstaller via : {% code overflow="wrap" expandable="true" %} \`\`\`bash uv venv unsloth-nvfp4-env --python 3.13 source unsloth-nvfp4-env/bin/activate uv pip install "vllm>=0.25.0" "flashinfer-python>=0.6.13" "nvidia-cutlass-dsl>=4.5.2" \\ --torch-backend=auto \`\`\` {% endcode %} Ensuite, pour servir dans vLLM pour DGX Spark : {% code overflow="wrap" expandable="true" %} \`\`\`shellscript export CUTE\_DSL\_ARCH=sm\_121a vllm serve unsloth/Qwen3.6-35B-A3B-NVFP4-Fast --moe-backend flashinfer\_b12x \`\`\` {% endcode %} Si vous rencontrez des problèmes avec Torchcodec, assurez-vous d’effectuer l’étape ci-dessous puis de relancer vllm. {% code overflow="wrap" expandable="true" %} \`\`\`bash sudo apt-get update sudo apt-get install -y ffmpeg \`\`\` {% endcode %} ### \*\*Tutoriel SGLang :\*\* Vous pouvez exécuter tous les modèles NVFP4 dans \[SGLang\](https://github.com/sgl-project/sglang). N’oubliez pas de remplacer le nom du modèle par celui que vous souhaitez. \*\*Qwen3.6 :\*\* \`\`\`bash python -m sglang.launch\_server --model-path unsloth/Qwen3.6-27B-NVFP4 --speculative-algorithm NEXTN \\ --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \`\`\` \*\*Gemma 4 :\*\* \`\`\`bash python -m sglang.launch\_server --model-path unsloth/Gemma-4-31B-NVFP4 --speculative-algorithm NEXTN \\ --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \`\`\` ### Gemma 4 et autres Chaque variante de Gemma 4 dispose désormais d’un checkpoint Unsloth Dynamic NVFP4. Nous montrons que Gemma-4 n’offre au maximum qu’un gain de débit de 1,44x en servant 128 personnes simultanément sur 1x B200 par rapport au BF16. Qwen3.5-122B-A10B est 1,38x plus rapide et GLM-4.7-Flash est 1,27x plus rapide. ![](https://unsloth.ai/files/5399d841bc5aca63b9229161857c2975dade37e5) \### Marlin vs Flashinfer vs cutlass vs cute-DSL Nous avons également constaté que les noyaux Marlin ne prennent pas bien en charge W4A4 - l’activer entraînera une dégradation des performances de 2,5x - utilisez donc CUTLASS, Flashinfer-TRTLLM ou Cute-DSL (activé automatiquement dans vLLM) ! De plus, si vous avez un DGX Spark, voir \[#dgx-spark-serving\](#dgx-spark-serving "mention") vous devez utiliser \`--moe-backend flashinfer\_b12x\` sinon, vous obtiendrez une inférence 2,5x plus lente. \*\*Donc, ne définissez aucun backend - vLLM sélectionne automatiquement le meilleur.\*\* | Modèle | schéma | backend | tok/s en décodage | débit tok/s | | --------------- | ------ | ------------------- | ----------------- | ----------- | | nvidia 27B | W4A16 | marlin (auto) | 115.6 | 2,403 | | unsloth 27B | W4A4 | marlin | 105.6 | 2,127 | | unsloth 27B | W4A4 | cutlass | 113.5 | 6,681 | | unsloth 27B | W4A4 | flashinfer\\\_trtllm | 112.6 | 6,158 | | unsloth 27B | W4A4 | \*\*cute-DSL (auto)\*\* | 125.9 | \*\*6,863\*\* | | nvidia 35B-A3B | W4A4 | marlin (auto) | 240.8 | 8,721 | | unsloth 35B-A3B | W4A4 | marlin | 215.8 | 8,619 | | unsloth 35B-A3B | W4A4 | cutlass | 158.3 | 11,017 | | unsloth 35B-A3B | W4A4 | \*\*cute-DSL (auto)\*\* | 295.2 | \*\*15,636\*\* | --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/nvfp4.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/kimi-k2.7-code.md). # Kimi K2.7 Code - Comment l'exécuter localement Kimi K2.7 Code est le modèle de codage agentique de Moonshot AI, s’appuyant sur \[K2.6\](/docs/fr/modeles/kimi-k2.6.md) pour améliorer l’achèvement des tâches tout en utilisant environ 30 % de jetons de réflexion en moins. Le modèle MoE de 1T de paramètres (32B actifs) prend en charge uniquement la réflexion, la vision et un contexte de 256K. Il offre des performances open SOTA sur les tâches de vision, de codage, agentiques, à long contexte et de chat. La précision complète nécessite 605 Go d’espace disque ; Unsloth \[Dynamique\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) Le 2 bits nécessite \*\*325 Go (-48 %)\*\*. Exécutez \[\*\*Kimi-K2.7-Code-GGUF\*\*\](https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF) via Unsloth Studio ou llama.cpp. \[\*\*Unsloth Dynamique\*\*\](/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs.md) \*\*quantifications\*\* met à niveau les couches importantes en 8 bits et nécessite \*\*310 Go+ VRAM/RAM\*\* configurations\*\*.\*\* Pour \*\*sans perte\*\* Pour Kimi K2.6, utilisez Q8 (\`UD-Q8\_K\_XL\`), qui n'est que \*\*10 Go de plus\*\* que Q4 (\`UD-Q4\_K\_XL\`). Vous pouvez exécuter Kimi K2.7 Code via un Mac Studio ou \[DGX Station\](/docs/fr/blog/dgx-station.md). \*\*Tableau : exigences matérielles\*\* (unités = mémoire totale : RAM + VRAM, ou mémoire unifiée) | 1 bit dynamique | Dynamique 2 bits | Q3 dynamique | Q8 (sans perte) | | --------------- | ---------------- | ------------ | --------------- | | 310 Go | 325-350 Go | 385-470 Go | 605 Go | ### 📊 Analyse de quantification Comme \[Kimi-K2.6\](/docs/fr/modeles/kimi-k2.6.md), \`UD-Q8\_K\_XL\` est sans perte car Kimi utilise int4 pour les poids MoE et BF16 pour tout le reste, et \`Q8\_K\_XL\` en découle. Ainsi, nous utilisons la même méthodologie dynamique pour la conversion de Kimi-K2.6. \`UD-Q4\_K\_XL\` est similaire, sauf que les tenseurs restants sont \`Q8\_0\`, donc il est presque en précision complète et nécessite 600 Go de RAM/VRAM. \`UD-Q8\_K\_XL\` est « véritablement sans perte ». | Mesure | UD-Q2\\\_K\\\_XL | UD-Q4\\\_K\\\_XL | UD-Q8\\\_K\\\_XL (sans perte) | | ------------- | ------------ | ------------ | ------------------------- | | Espace disque | 339 Go | 584 Go | 595 Go | | Perplexité | \\~2.4131 | \\~1.8420 | \\~1.8419 | Nous avons suivi \[jukofyork\](https://github.com/jukofyork)a découvert que \`const float d = max / -7;\` au lieu de la valeur par défaut \`const float d = max / -8;\` pendant le processus de quantification uniquement sur les couches MoE. Ce correctif de bijection sur les MoE natives INT4 permet au \`Q4\_0\` type de quantification de réduire l'erreur absolue de 1,8 % à presque 0 % (epsilon). Par exemple, ci-dessous se trouve l'histogramme pour Kimi-K2.7-Code, et vous pouvez voir que -8 n'est absolument pas utilisé : ![](https://unsloth.ai/files/7b976d647b9e4d9a81c4a292e28fbc45ac3ddcc6) Notez que nous devons également conserver les autres couches en BF16 et ne pas utiliser le « Q4\\\_0 » intelligent. Nous montrons ci-dessous les graphiques d'erreur pour les deux par rapport à la base BF16. \`UD-Q8-K\_XL\` est vraiment « sans perte », avec une différence de l'ordre de l'epsilon machine lors de la conversion de Q4\\\_0 vers BF16. Ainsi, Q4\\\_K\\\_XL présente bien une certaine erreur de quantification due à l'utilisation de Q8\\\_0, tandis que Q8\\\_K\\\_XL est presque sans perte, sauf pour l'arrondi BF16. ![](https://unsloth.ai/files/b96158e5a2dd307aa5a0ff4002cf1098e22dcd84) Pour Q4\\\_K\\\_XL, nous traçons également l'erreur par tenseur de Q8\\\_0 par rapport à BF16. En général, il existe une certaine erreur entre Q8\\\_K\\\_XL (presque sans perte) et Q4\\\_K\\\_XL, mais pas beaucoup. ![](https://unsloth.ai/files/e7b64c71af3efca3dfcffae959133e227389eaed) \### :gear: Guide d'utilisation Kimi K2.7 Code est \*\*réflexion uniquement\*\*, avec \*\*\`preserve\_thinking\` toujours activé\*\*. Le mode instantané n’est pas pris en charge. | Par défaut (mode réflexion) | | --------------------------- | | temperature = 1.0 | | top\\\_p = 0,95 | \* Longueur de contexte suggérée = \`98,304\` (jusqu’à \`262,144\`) Si le modèle tient en mémoire, vous obtiendrez >100 jetons/s avec des B200. Nous recommandons \`UD-Q2\_K\_XL\` (345 Go) comme bon compromis taille/qualité. Meilleure règle empirique : RAM+VRAM ≈ la taille de la quantification ; sinon, cela fonctionnera quand même, mais plus lentement à cause du déchargement. #### Gabarit de chat pour Kimi K2.7-Code Exécution \`tokenizer.apply\_chat\_template(\[{"role": "user", "content": "Combien font 1+1 ?"},\])\` donne : {% code overflow="wrap" %} \`\`\` <|im\_user|>user<|im\_middle|>Combien font 1+1 ?<|im\_end|><|im\_assistant|>assistant<|im\_middle|> \`\`\` {% endcode %} Si nous saisissons également les outils comme référencés dans \[Tool Calling Guide\](/docs/fr/notions-de-base/tool-calling-guide-for-local-llms.md), alors voici ce que l’on obtient : {% code overflow="wrap" expandable="true" %} \`\`\` <|im\_system|>tool\_declare<|im\_middle|># Outils ## fonctions namespace functions { // Additionner deux nombres. type add\_number = (\_: { // Le premier nombre. a: string, // Le deuxième nombre. b: string }) => any; // Multiplier deux nombres. type multiply\_number = (\_: { // Le premier nombre. a: string, // Le deuxième nombre. b: string }) => any; // Soustraire deux nombres. type subtract\_number = (\_: { // Le premier nombre. a: string, // Le deuxième nombre. b: string }) => any; // Écrire une histoire aléatoire. type write\_a\_story = (\_: {}) => any; // Effectuer des opérations dans le terminal. type terminal = (\_: { // La commande que vous souhaitez lancer, par ex. \`ls\`, \`rm\`, ... command: string }) => any; // Appeler un interpréteur Python avec du code Python qui sera exécuté. type python = (\_: { // Le code Python à exécuter code: string }) => any; } <|im\_end|><|im\_user|>user<|im\_middle|>Combien font 1+1 ?<|im\_end|><|im\_assistant|>assistant<|im\_middle|> \`\`\` {% endcode %} ## Guide d’exécution de Kimi K2.7 Code ### 🦥 Exécuter Kimi-K2.7-Code dans Unsloth Studio Kimi K2.7 Code peut s’exécuter dans \[Unsloth Studio\](/docs/fr/nouveau/studio.md), une interface web open source pour l’IA locale. \*\*Unsloth Studio décharge automatiquement vers la RAM et détecte les configurations multi-GPU\*\*. Avec Unsloth Studio, vous pouvez exécuter des modèles localement sur \*\*MacOS, Windows\*\*, Linux et : {% columns %} {% column %} \* Rechercher, télécharger, \[exécuter des GGUF\](/docs/fr/nouveau/studio.md#run-models-locally) et des modèles safetensor \* \[\*\*Appels d'outils auto-réparateurs\*\* appels d'outils\](/docs/fr/nouveau/studio.md#execute-code--heal-tool-calling) + \*\*recherche web\*\* \* \[\*\*Exécution de code\*\*\](/docs/fr/nouveau/studio.md#run-models-locally) (Python, Bash) \* \[Inférence automatique\](/docs/fr/nouveau/studio.md#model-arena) réglage des paramètres (temp, top-p, etc.) \* Inférence rapide CPU + GPU via llama.cpp \* \[Entraîner des LLM\](/docs/fr/nouveau/studio.md#no-code-training) 2x plus vite avec 70 % de VRAM en moins {% endcolumn %} {% column %} ![](https://unsloth.ai/files/9d149ac4b773a56a635d40ab8347ea2896781ae6) {% endcolumn %} {% endcolumns %} {% stepper %} {% step %} \*\*Installer et lancer Unsloth\*\* Pour l’installer, exécutez dans votre terminal : MacOS, Linux, WSL : \`\`\`bash curl -fsSL https://unsloth.ai/install.sh | sh \`\`\` Windows PowerShell : \`\`\`bash irm https://unsloth.ai/install.ps1 | iex \`\`\` \*\*Lancer Unsloth\*\* MacOS, Linux, WSL et Windows : \`\`\`bash unsloth studio -H 0.0.0.0 -p 8888 \`\`\` Puis ouvrez \`http://127.0.0.1:8888\` (ou votre URL spécifique) dans votre navigateur. {% endstep %} {% step %} \*\*Recherchez et téléchargez Kimi K2.7-Code\*\* Unsloth Studio décharge automatiquement vers la RAM et détecte les configurations multi-GPU. Lors du premier lancement, vous devrez créer un mot de passe pour sécuriser votre compte et vous reconnecter plus tard. Puis allez dans l’ \[Unsloth Chat\](/docs/fr/nouveau/studio/chat.md) onglet et recherchez \*\*Kimi-K2.7 Code\*\* dans la barre de recherche et téléchargez le modèle et la quantification souhaités. Assurez-vous d’avoir suffisamment de puissance de calcul pour exécuter le modèle. ![](https://unsloth.ai/files/1255891bf200c1f40c2d752f916601f73e88c543) {% endstep %} {% step %} \*\*Exécuter Kimi-K2.7-Code\*\* Les paramètres d’inférence devraient être définis automatiquement lors de l’utilisation d’Unsloth Studio, mais vous pouvez toujours les modifier manuellement. Vous pouvez également modifier la longueur du contexte, le modèle de chat et d’autres paramètres. Pour plus d'informations, vous pouvez consulter notre \[guide d'inférence Unsloth Studio\](/docs/fr/nouveau/studio/chat.md). ![](https://unsloth.ai/files/acbddf2b13951f7e83eb5a222093fd827a4b5258) Exemple de Qwen3.6 fonctionnant avec appel d’outils {% endstep %} {% endstepper %} ### 🦙 Exécuter Kimi K2.7 Code dans llama.cpp Pour ce guide, nous allons exécuter le \`UD-Q2\_K\_XL\` quant qui nécessitera au moins 345 Go de RAM. N’hésitez pas à modifier le type de quantification. GGUF : \[\*\*Kimi-K2.7-Code-GGUF\*\*\](https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF) Pour ces tutoriels, nous utiliserons \[llama.cpp\](llama.cpphttps://github.com/ggml-org/llama.cpp) pour une inférence locale rapide, surtout si vous avez un CPU. {% stepper %} {% step %} Obtenez la dernière \`llama.cpp\` \*\*sur\*\* \[\*\*GitHub ici\*\*\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez \`-DGGML\_CUDA=ON\` à \`-DGGML\_CUDA=OFF\` si vous n'avez pas de GPU ou si vous voulez simplement une inférence CPU. \*\*Pour les appareils Apple Mac / Metal\*\*, définissez \`-DGGML\_CUDA=OFF\` puis continuez comme d'habitude - la prise en charge de Metal est activée par défaut. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\\\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` {% endstep %} {% step %} \*\*Prenons d’abord une image !\*\* Vous pouvez également téléverser des images. Nous utiliserons , qui n’est que notre mini-logo montrant comment les fine-tunes sont réalisés avec Unsloth : {% code overflow="wrap" %} \`\`\`bash wget https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/unsloth%20made%20with%20love.png -O unsloth.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/8226e799b1a697c94669626a5da8371ee8388470) Prenons aussi la 2e image à {% code overflow="wrap" %} \`\`\`bash wget https://files.worldwildlife.org/wwfcmsprod/images/Sloth\_Sitting\_iStock\_3\_12\_2014/story\_full\_width/8l7pbjmj29\_iStock\_000011145477Large\_mini\_\_1\_.jpg -O picture.png \`\`\` {% endcode %} ![](https://unsloth.ai/files/eb404a2bbe8e7f1836a2608a3950391d566357e3) {% endstep %} {% step %} Vous pouvez maintenant utiliser \`llama.cpp\` directement pour charger et télécharger des modèles, tout comme \`ollama run\`. Tout d'abord, sélectionnez le type de quantification que vous souhaitez, comme \`Q2\_K\_XL\`. Utilisez aussi \`export LLAMA\_CACHE="folder"\` pour forcer \`llama.cpp\` l'enregistrement dans un emplacement spécifique. Notez que ce processus de téléchargement peut être très lent, il est donc probablement préférable d'utiliser le processus de téléchargement manuel dans la section suivante. \`\`\`bash export LLAMA\_CACHE="unsloth/Kimi-K2.7-Code-GGUF" ./llama.cpp/llama-cli \\\\ -hf unsloth/Kimi-K2.7-Code-GGUF:UD-Q2\_K\_XL \\\\ --temp 1.0 \\ --top-p 0.95 \`\`\` {% endstep %} {% step %} Si vous souhaitez télécharger le modèle manuellement, nous pouvons le télécharger via le code ci-dessous (après avoir installé \`pip install huggingface\_hub\`). Si les téléchargements restent bloqués, voir : \[Hugging Face Hub, débogage XET\](/docs/fr/notions-de-base/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md) \`\`\`bash hf download unsloth/Kimi-K2.7-Code-GGUF \\\\ --local-dir unsloth/Kimi-K2.7-Code-GGUF \\\\ --include "\*mmproj-F16\*" \\\\ --include "\*UD-Q2\_K\_XL\*" # Utilisez "\*UD-Q8\_K\_XL\*" pour la précision complète \`\`\` {% endstep %} {% step %} Puis exécutez le modèle en mode conversation : {% code overflow="wrap" %} \`\`\`bash ./llama.cpp/llama-cli \\\\ --model unsloth/Kimi-K2.7-Code-GGUF/UD-Q2\_K\_XL/Kimi-K2.7-Code-UD-Q2\_K\_XL-00001-of-00008.gguf \\\\ --mmproj unsloth/Kimi-K2.7-Code-GGUF/mmproj-F16.gguf \\\\ --temp 1.0 \\ --top-p 0.95 \`\`\` {% endcode %} Vous verrez alors ce qui suit :\\ !\[\](/files/e6af1d459bcfdf53b2e2457b268ac8b1ba9cce95) {% endstep %} {% step %} Ensuite, utilisez \`/image\` pour charger les deux images et demander « Quelle est cette image » : ![](https://unsloth.ai/files/293e043d360bafc05607d13d403ce69ebd8b5bb8) et vous obtiendrez quelque chose comme ci-dessous : ![](https://unsloth.ai/files/222b0cb3d6f4bff9098a9b3997e6f2d77e619237) Sur la 2e image du paresseux : ![](https://unsloth.ai/files/0fbf40a1fb039b839019a79dcf7367eea9ca95bf) Ce qui vous donnera : ![](https://unsloth.ai/files/0d381287e2e61e78d5f5c9d3bca19555fdab0baf) {% endstep %} {% endstepper %} ### 📊 Benchmarks Vous pouvez consulter ci-dessous d’autres benchmarks sous forme de tableau : ![](https://unsloth.ai/files/9722c6772805327e152980b803bf3a4a22879947) | Benchmark | Kimi K2.7 Code | Kimi K2.6 | GPT-5.5 | Claude Opus 4.8 | | :-------------------------: | :------------: | :-------: | :-----: | :-------------: | | \*\*Codage\*\* | | | | | | Kimi Code Bench v2 | 62.0 | 50.9 | 69.0 | 67.4 | | Banc d’essai de programme | 53.6 | 48.3 | 69.1 | 63.8 | | MLS Bench Lite | 35.1 | 26.7 | 35.5 | 42.8 | | \*\*Agentique\*\* | | | | | | Banc d’essai Kimi Claw 24/7 | 46.9 | 42.9 | 52.8 | 50.4 | | MCP Atlas | 76.0 | 69.4 | 79.4 | 81.3 | | MCP Mark vérifié | 81.1 | 72.8 | 92.9 | 76.4 | --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/kimi-k2.7-code.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/notions-de-base/text-to-speech-tts-fine-tuning.md). # Guide de fine-tuning Text-to-Speech (TTS) L’affinage des modèles TTS leur permet de s’adapter à votre jeu de données spécifique, à votre cas d’usage, ou au style et au ton souhaités. L’objectif est de personnaliser ces modèles pour cloner des voix, adapter des styles et des tons de parole, prendre en charge de nouvelles langues, gérer des tâches spécifiques, et plus encore. Nous prenons également en charge \*\*Speech-to-Text (STT)\*\* des modèles comme Whisper d’OpenAI. Avec \[Unsloth\](https://github.com/unslothai/unsloth), vous pouvez affiner \*\*n’importe quel\*\* modèle TTS (\`transformers\` compatible) 1,5× plus rapidement avec 50 % de mémoire en moins que d’autres implémentations avec Flash Attention 2. ⭐ \*\*Unsloth prend en charge n’importe quel \`transformers\` modèle TTS compatible.\*\* Même si nous n’avons pas encore de notebook ou de téléversement pour celui-ci, il est tout de même pris en charge, par exemple essayez d’affiner Dia-TTS ou Moshi. {% hint style="info" %} Le clonage zero-shot capture le ton mais rate le rythme et l’expression, donnant souvent un rendu robotique et artificiel. L’affinage fournit une reproduction vocale bien plus précise et réaliste. \[En savoir plus ici\](#fine-tuning-voice-models-vs.-zero-shot-voice-cloning). {% endhint %} ### Notebooks d’affinage : Nous avons également téléversé des modèles TTS (originaux et quantifiés) sur notre \[page Hugging Face\](https://huggingface.co/collections/unsloth/text-to-speech-tts-models-68007ab12522e96be1e02155). | \[Sesame-CSM (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Sesame\_CSM\_\\(1B\\)-TTS.ipynb) | \[Orpheus-TTS (3B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Orpheus\_\\(3B\\)-TTS.ipynb) | \[Whisper Large V3\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Whisper.ipynb) (STT) | | ------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | \[Spark-TTS (0,5B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Spark\_TTS\_\\(0\_5B\\).ipynb) | \[Llasa-TTS (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llasa\_TTS\_\\(1B\\).ipynb) | \[Oute-TTS (1B)\](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Oute\_TTS\_\\(1B\\).ipynb) | {% hint style="success" %} Si vous remarquez que la durée de sortie atteint un maximum de 10 secondes, augmentez\`max\_new\_tokens = 125\` par rapport à sa valeur par défaut de 125. Comme 125 jetons correspondent à 10 secondes d’audio, vous devrez définir une valeur plus élevée pour des sorties plus longues. {% endhint %} ### Choisir et charger un modèle TTS Pour le TTS, les modèles plus petits sont souvent préférés en raison d’une latence plus faible et d’une inférence plus rapide pour les utilisateurs finaux. Affiner un modèle de moins de 3 milliards de paramètres est souvent idéal, et nos principaux exemples utilisent Sesame-CSM (1B) et Orpheus-TTS (3B), un modèle vocal basé sur Llama. #### Détails de Sesame-CSM (1B) \*\*CSM-1B\*\* est un modèle de base, tandis que \*\*Orpheus-ft\*\* est affiné sur 8 comédiens professionnels, ce qui fait de la cohérence vocale la principale différence. CSM nécessite un contexte audio pour chaque locuteur afin d’obtenir de bonnes performances, tandis qu’Orpheus-ft intègre cette cohérence nativement. L’affinage à partir d’un modèle de base comme CSM nécessite généralement plus de calcul, tandis que partir d’un modèle déjà affiné comme Orpheus-ft offre de meilleurs résultats dès le départ. Pour aider avec CSM, nous avons ajouté de nouvelles options d’échantillonnage et un exemple montrant comment utiliser le contexte audio pour améliorer la cohérence vocale. #### Détails d’Orpheus-TTS (3B) Orpheus est préentraîné sur un vaste corpus vocal et excelle dans la génération d’une parole réaliste avec une prise en charge intégrée d’indices émotionnels comme les rires et les soupirs. Son architecture en fait l’un des modèles TTS les plus faciles à utiliser et à entraîner, car il peut être exporté via llama.cpp, ce qui lui confère une excellente compatibilité avec tous les moteurs d’inférence. Pour les modèles non pris en charge, vous ne pourrez enregistrer que les safetensors de l’adaptateur LoRA. #### Chargement des modèles Comme les modèles vocaux sont généralement de petite taille, vous pouvez les entraîner en utilisant LoRA 16 bits ou un affinage complet FFT, ce qui peut fournir des résultats de meilleure qualité. Pour le charger en LoRA 16 bits : \`\`\`python from unsloth import FastModel model\_name = "unsloth/orpheus-3b-0.1-pretrained" model, tokenizer = FastModel.from\_pretrained( model\_name, load\_in\_4bit=False # utiliser une précision 4 bits (QLoRA) ) \`\`\` Lorsque cela s’exécute, Unsloth téléchargera les poids du modèle ; si vous préférez 8 bits, vous pouvez utiliser \`load\_in\_8bit = True\`, ou pour un affinage complet définissez \`full\_finetuning = True\` (assurez-vous d’avoir suffisamment de VRAM). Vous pouvez également remplacer le nom du modèle par d’autres modèles TTS. {% hint style="info" %} \*\*Remarque :\*\* Le tokenizer d’Orpheus inclut déjà des jetons spéciaux pour la sortie audio (plus d’informations à ce sujet plus tard). Vous n’avez \*pas\* besoin d’un vocodeur séparé – Orpheus produira directement des jetons audio, qui pourront être décodés en forme d’onde. {% endhint %} ### Préparer votre jeu de données Au minimum, un jeu de données pour l’affinage TTS se compose de \*\*clips audio et de leurs transcriptions correspondantes\*\* (texte). Utilisons le \[\*Elise\* jeu de données\](https://huggingface.co/datasets/MrDragonFox/Elise) qui est un corpus de parole anglaise d’environ 3 heures, avec un seul locuteur. Il existe deux variantes : \* \[\`MrDragonFox/Elise\`\](https://huggingface.co/datasets/MrDragonFox/Elise) – une version augmentée avec des \*\*étiquettes d’émotion\*\* (par ex. \\, \\) intégrées dans les transcriptions. Ces étiquettes entre chevrons indiquent des expressions (rire, soupirs, etc.) et sont traitées comme des jetons spéciaux par le tokenizer d’Orpheus \* \[\`Jinsaryko/Elise\`\](https://huggingface.co/datasets/Jinsaryko/Elise) – version de base avec des transcriptions sans étiquettes spéciales. Le jeu de données est organisé avec un audio et une transcription par entrée. Sur Hugging Face, ces jeux de données ont des champs tels que \`audio\` (la forme d’onde), \`text\` (la transcription), ainsi que કેટલીક métadonnées (nom du locuteur, statistiques de hauteur, etc.). Nous devons fournir à Unsloth un jeu de données de paires audio-texte. {% hint style="success" %} Plutôt que de se concentrer uniquement sur le ton, la cadence et la hauteur, la priorité devrait être de s’assurer que votre jeu de données est entièrement annoté et correctement normalisé. {% endhint %} {% hint style="info" %} Avec certains modèles comme \*\*Sesame-CSM-1B\*\*, vous pourriez remarquer une variation de voix entre les générations en utilisant l’ID de locuteur 0 parce que c’est un \*\*modèle de base\*\*— il n’a pas d’identités vocales fixes. Les jetons d’ID de locuteur aident principalement à maintenir \*\*la cohérence au sein d’une conversation\*\*, pas entre des générations séparées. Pour obtenir une voix cohérente, fournissez des \*\*exemples contextuels\*\*, comme quelques extraits audio de référence ou des répliques précédentes. Cela aide le modèle à imiter la voix souhaitée de manière plus fiable. Sans cela, une variation est normale, même avec le même ID de locuteur. {% endhint %} \*\*Option 1 : utiliser la bibliothèque Hugging Face Datasets\*\* – Nous pouvons charger le jeu de données Elise à l’aide de la \`bibliothèque datasets\` de Hugging Face : \`\`\`python from datasets import load\_dataset, Audio # Charger le jeu de données Elise (par exemple, la version avec étiquettes d’émotion) dataset = load\_dataset("MrDragonFox/Elise", split="train") print(len(dataset), "exemples") # ~1200 exemples dans Elise # S’assurer que tout l’audio est à un taux d’échantillonnage de 24 kHz (taux attendu par Orpheus) dataset = dataset.cast\_column("audio", Audio(sampling\_rate=24000)) \`\`\` Cela téléchargera le jeu de données (\\~328 Mo pour \\~1,2k exemples). Chaque élément de \`jeu de données\` est un dictionnaire contenant au moins : \* \`"audio"\`: le clip audio (tableau de forme d’onde et métadonnées comme le taux d’échantillonnage), et \* \`"text"\`: la chaîne de transcription Orpheus prend en charge des étiquettes comme \`\`, \`\`, \`\`, \`\`, \`\`, \`\`, \`\`, \`\`, etc. Par exemple : \`"I missed you so much!"\`. Ces étiquettes sont entre chevrons et seront traitées comme des jetons spéciaux par le modèle (elles correspondent aux \[étiquettes attendues par Orpheus\](https://github.com/canopyai/Orpheus-TTS) comme \`\` et \`\`. Pendant l’entraînement, le modèle apprendra à associer ces étiquettes aux motifs audio correspondants. Le jeu de données Elise avec étiquettes contient déjà beaucoup d’entre elles (par exemple, 336 occurrences de « laughs », 156 de « sighs », etc., comme indiqué dans sa fiche). Si votre jeu de données n’a pas de telles étiquettes mais que vous souhaitez les intégrer, vous pouvez annoter manuellement les transcriptions lorsque l’audio contient ces expressions. \*\*Option 2 : préparer un jeu de données personnalisé\*\* – Si vous avez vos propres fichiers audio et transcriptions : \* Organisez les clips audio (fichiers WAV/FLAC) dans un dossier. \* Créez un fichier CSV ou TSV avec des colonnes pour le chemin du fichier et la transcription. Par exemple : \`\`\` filename,text 0001.wav,Bonjour ! 0002.wav, Je suis très fatigué. \`\`\` \* Utilisez \`load\_dataset("csv", data\_files="mydata.csv", split="train")\` pour le charger. Vous devrez peut-être indiquer au chargeur de jeu de données comment gérer les chemins audio. Une autre option consiste à utiliser la fonctionnalité \`datasets.Audio\` pour charger les données audio à la volée : \`\`\`python from datasets import Audio dataset = load\_dataset("csv", data\_files="mydata.csv", split="train") dataset = dataset.cast\_column("filename", Audio(sampling\_rate=24000)) \`\`\` Ensuite \`dataset\[i\]\["audio"\]\` contiendra le tableau audio. \* \*\*Assurez-vous que les transcriptions sont normalisées\*\* (pas de caractères inhabituels que le tokenizer pourrait ne pas connaître, sauf les étiquettes d’émotion si elles sont utilisées). Assurez-vous également que tout l’audio a un taux d’échantillonnage cohérent (rééchantillonnez-le si nécessaire au taux cible attendu par le modèle, par ex. 24 kHz pour Orpheus). En résumé, pour \*\*la préparation du jeu de données\*\*: \* vous avez besoin d’une \*\*liste de paires (audio, texte)\*\* . \* Utilisez la \`bibliothèque datasets\` bibliothèque HF \* pour gérer le chargement et le prétraitement optionnel (comme le rééchantillonnage). \*\*Incluez toutes les\*\* étiquettes spéciales \`dans le texte que vous souhaitez faire apprendre au modèle (assurez-vous qu’elles sont au format\` \\ \* afin que le modèle les traite comme des jetons distincts). ### (Facultatif) Si vous avez plusieurs locuteurs, vous pourriez inclure un jeton d’ID de locuteur dans le texte ou utiliser une approche séparée d’empreinte vocale, mais cela dépasse ce guide de base (Elise est monolocuteur). Affinage de TTS avec Unsloth \*\*Commençons maintenant l’affinage ! Nous allons illustrer cela en utilisant du code Python (que vous pouvez exécuter dans un notebook Jupyter, Colab, etc.).\*\* Étape 1 : charger le modèle et le jeu de données \`Dans tous nos notebooks TTS, nous activons l’entraînement LoRA (16 bits) et désactivons l’entraînement QLoRA (4 bits) avec :\`load\\\_in\\\_4bit = False \`\`\`python . Cela permet généralement au modèle d’apprendre mieux votre jeu de données et d’obtenir une plus grande précision. from unsloth import FastLanguageModel import torch dtype = None # None pour la détection automatique. Float16 pour Tesla T4, V100, Bfloat16 pour Ampere+ load\_in\_4bit = False # Utiliser la quantification 4 bits pour réduire l’utilisation mémoire. Peut être False. model, tokenizer = FastLanguageModel.from\_pretrained( model\_name = "unsloth/orpheus-3b-0.1-ft", max\_seq\_length= 2048, # Choisissez n’importe quelle valeur pour un long contexte ! dtype = dtype, load\_in\_4bit = load\_in\_4bit, ) #token = "hf\_...", # utilisez-en un si vous utilisez des modèles soumis à restriction comme meta-llama/Llama-2-7b-hf from datasets import load\_dataset \`\`\` {% hint style="info" %} dataset = load\\\_dataset("MrDragonFox/Elise", split = "train") {% endhint %} \*\*Si la mémoire est très limitée ou si le jeu de données est volumineux, vous pouvez diffuser ou charger par morceaux. Ici, 3 h d’audio tiennent facilement en RAM. Si vous utilisez votre propre CSV de jeu de données, chargez-le de la même manière.\*\* Étape 2 : avancé - prétraiter les données pour l’entraînement (facultatif) \`\`\`python Nous devons préparer les entrées pour le Trainer. Pour la synthèse vocale, une approche consiste à entraîner le modèle de manière causale : concaténer le texte et les IDs des jetons audio comme séquence cible. Cependant, comme Orpheus est un LLM décodeur uniquement qui produit de l’audio, nous pouvons fournir le texte en entrée (contexte) et utiliser les IDs des jetons audio comme étiquettes. En pratique, l’intégration d’Unsloth peut faire cela automatiquement si la configuration du modèle l’identifie comme un modèle de text-to-speech. Si ce n’est pas le cas, nous pouvons faire quelque chose comme : # Tokeniser les transcriptions textuelles def preprocess\_function(example): # Tokeniser le texte (conserver intacts les jetons spéciaux comme ) tokens = tokenizer(example\["text"\], return\_tensors="pt") # Aplatir en liste d’IDs de jetons input\_ids = tokens\["input\_ids"\].squeeze(0) # Le modèle générera des jetons audio après ces jetons texte. # Pour l’entraînement, nous pouvons définir labels égaux à input\_ids (afin qu’il apprenne à prédire le jeton suivant). # Mais cela ne couvre que les jetons texte prédisant le jeton texte suivant (qui peut être un jeton audio ou une fin). # Une approche plus sophistiquée : ajouter un jeton spécial indiquant le début de l’audio, et laisser le modèle générer le reste. # Pour simplifier, utilisez la même entrée que les labels (le modèle apprendra à produire la séquence en se basant sur elle-même). return {"input\_ids": input\_ids, "labels": input\_ids} \`\`\` {% hint style="info" %} train\\\_data = dataset.map(preprocess\\\_function, remove\\\_columns=dataset.column\\\_names) \*Ce qui précède est une simplification. En réalité, pour affiner Orpheus correctement, vous auriez besoin des\*jetons audio comme partie des labels d’entraînement \`. Le préentraînement d’Orpheus impliquait probablement la conversion de l’audio en jetons discrets (via un codec audio) et l’entraînement du modèle à les prédire à partir du texte précédent. Pour un affinage sur de nouvelles données vocales, vous devriez de la même manière obtenir les jetons audio pour chaque clip (en utilisant le codec audio d’Orpheus). Le GitHub d’Orpheus fournit un script de traitement des données – il encode l’audio en séquences de\` \\ {% endhint %} jetons. \*\*Cependant,\*\*Unsloth peut abstraire cela : si le modèle est un FastModel avec un processeur associé qui sait gérer l’audio, il pourrait encoder automatiquement l’audio du jeu de données en jetons. Sinon, vous devrez encoder manuellement chaque clip audio en IDs de jetons (en utilisant le codebook d’Orpheus). C’est une étape avancée au-delà de ce guide, mais gardez à l’esprit que le simple fait d’utiliser des jetons texte n’enseignera pas au modèle l’audio réel – il doit correspondre aux motifs audio. \`processor\` et en passant le tableau audio). Si Unsloth ne prend pas encore en charge la tokenisation audio automatique, vous devrez peut-être utiliser la fonction \`encode\_audio\` du dépôt Orpheus pour obtenir les séquences de jetons de l’audio, puis les utiliser comme labels. (Les entrées du jeu de données ont bien des \`phonèmes\` et certaines caractéristiques acoustiques, ce qui suggère un pipeline.) \*\*Étape 3 : configurer les arguments d’entraînement et le Trainer\*\* \`\`\`python from transformers import TrainingArguments,Trainer,DataCollatorForSeq2Seq from unsloth import is\_bfloat16\_supported trainer = Trainer( model = model, train\_dataset = dataset, args = TrainingArguments( per\_device\_train\_batch\_size = 1, gradient\_accumulation\_steps = 4, warmup\_steps = 5, # num\_train\_epochs = 1, # Définissez ceci pour une exécution d’entraînement complète. max\_steps = 60, learning\_rate = 2e-4, fp16 = not is\_bfloat16\_supported(), bf16 = is\_bfloat16\_supported(), logging\_steps = 1, optim = "adamw\_8bit", weight\_decay = 0.01, lr\_scheduler\_type = "linear", seed = 3407, output\_dir = "outputs", report\_to = "none", # Utilisez ceci pour WandB, etc. ), ) \`\`\` Nous faisons 60 étapes pour accélérer les choses, mais vous pouvez définir \`num\_train\_epochs=1\` pour une exécution complète, et désactiver \`max\_steps=None\`. Utiliser per\\\_device\\\_train\\\_batch\\\_size >1 peut entraîner des erreurs si une configuration multi-GPU est utilisée ; pour éviter les problèmes, assurez-vous que CUDA\\\_VISIBLE\\\_DEVICES est défini sur un seul GPU (par ex. CUDA\\\_VISIBLE\\\_DEVICES=0). Ajustez selon vos besoins. \*\*Étape 4 : commencer l’affinage\*\* Cela lancera la boucle d’entraînement. Vous devriez voir les journaux de perte toutes les 50 étapes (comme défini par \`logging\_steps\`). L’entraînement peut prendre un certain temps selon le GPU – par exemple, sur un GPU Colab T4, quelques époques sur 3 h de données peuvent prendre 1 à 2 heures. Les optimisations d’Unsloth le rendront plus rapide que l’entraînement HF standard. \*\*Étape 5 : enregistrer le modèle affiné\*\* Une fois l’entraînement terminé (ou si vous l’arrêtez en cours de route lorsque vous estimez que c’est suffisant), enregistrez le modèle. Cela enregistre UNIQUEMENT les adaptateurs LoRA, et non le modèle complet. Pour enregistrer en 16 bits ou en GGUF, faites défiler vers le bas ! \`\`\`python model.save\_pretrained("lora\_model") # Enregistrement local tokenizer.save\_pretrained("lora\_model") # model.push\_to\_hub("your\_name/lora\_model", token = "...") # Enregistrement en ligne # tokenizer.push\_to\_hub("your\_name/lora\_model", token = "...") # Enregistrement en ligne \`\`\` Cela enregistre les poids du modèle (pour LoRA, il peut n’enregistrer que les poids de l’adaptateur si la base n’est pas entièrement affinée). Si vous avez utilisé \`--push\_model\` dans la CLI ou \`trainer.push\_to\_hub()\`, vous pourriez le téléverser directement sur Hugging Face Hub. Vous devriez maintenant avoir un modèle TTS affiné dans le répertoire. L’étape suivante consiste à le tester et, si pris en charge, vous pouvez utiliser llama.cpp pour le convertir en fichier GGUF. ### Affinage des modèles vocaux vs clonage vocal zero-shot On dit qu’on peut cloner une voix avec seulement 30 secondes d’audio en utilisant des modèles comme XTTS — sans entraînement nécessaire. C’est techniquement vrai, mais cela passe à côté de l’essentiel. Le clonage vocal zero-shot, également disponible dans des modèles comme Orpheus et CSM, est une approximation. Il capture le \*\*ton et le timbre\*\* généraux de la voix d’un locuteur, mais ne reproduit pas toute la gamme expressive. Vous perdez des détails comme la vitesse de parole, la formulation, les particularités vocales et les subtilités de la prosodie — des éléments qui donnent à une voix sa \*\*personnalité et son unicité\*\*. Si vous voulez juste une voix différente et que le même mode d’élocution vous convient, le zero-shot est généralement suffisant. Mais la parole suivra toujours le \*\*style du modèle\*\*, et non celui du locuteur. Pour quelque chose de plus personnalisé ou expressif, vous avez besoin d’un entraînement avec des méthodes comme LoRA pour vraiment capturer la manière dont quelqu’un parle. --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/notions-de-base/text-to-speech-tts-fine-tuning.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Unknown \> For the complete documentation index, see \[llms.txt\](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending \`.md\` to page URLs; this page is available as \[Markdown\](https://unsloth.ai/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md). # Entraînement gpt-oss à long contexte Nous sommes ravis de présenter la prise en charge d’Unsloth Flex Attention pour l’entraînement OpenAI gpt-oss, ce qui permet \*\*>8× plus longues longueurs de contexte\*\*, \*\*>50 % d’utilisation de VRAM en moins\*\* et \*\*un entraînement >1,5× plus rapide (sans dégradation de la précision)\*\* par rapport à toutes les implémentations, y compris celles utilisant Flash Attention 3 (FA3). Unsloth Flex Attention permet de s’entraîner avec une \*\*longueur de contexte de 60K\*\* sur un GPU H100 avec 80 Go de VRAM pour BF16 LoRA. Aussi : \* Vous pouvez \[désormais exporter/enregistrer\](#new-saving-to-gguf-vllm-after-gpt-oss-training) votre modèle gpt-oss affiné en QLoRA vers llama.cpp, vLLM, Ollama ou HF \* Nous \[\*\*avons corrigé l’entraînement gpt-oss\*\*\](#bug-fixes-for-gpt-oss) \*\*les pertes qui divergeaient vers l’infini\*\* sur des GPU float16 (comme le T4 Colab) \* Nous \[avons corrigé l’implémentation gpt-oss\](#bug-fixes-for-gpt-oss) des problèmes sans rapport avec Unsloth, notamment en garantissant que \`swiglu\_limit = 7.0\` est correctement appliqué lors de l’inférence MXFP4 dans transformers ## 🦥Présentation de la prise en charge d’Unsloth Flex Attention Avec la prise en charge de Flex Attention par Unsloth, un seul H100 avec 80 Go de VRAM peut gérer jusqu’à 81K de longueur de contexte avec QLoRA et 60K de contexte avec BF16 LoRA ! Ces gains s’appliquent à \*\*LES DEUX\*\* gpt-oss-20b et \*\*gpt-oss-120b\*\*! Plus vous utilisez une grande longueur de contexte, plus vous gagnerez avec Unsloth Flex Attention : ![](https://unsloth.ai/files/eeb537df86214dd9cc6b3e73b024ee8299a4bbac) En comparaison, toutes les autres implémentations non-Unsloth plafonnent à 9K de longueur de contexte sur un GPU de 80 Go, et ne peuvent atteindre que 15K de contexte avec FA3. Mais, \*\*FA3 ne convient pas à l’entraînement de gpt-oss car il ne prend pas en charge la rétropropagation pour les puits d’attention\*\*. Donc, si vous utilisiez auparavant FA3 pour l’entraînement de gpt-oss, nous vous recommandons de \*\*ne pas l’utiliser\*\* pour le moment. Ainsi, la longueur de contexte maximale que vous pouvez obtenir sans Unsloth sur 80 Go de VRAM est d’environ 9K. L’entraînement avec Unsloth Flex Attention offre au moins un accélération de 1,3×, avec des gains qui augmentent à mesure que la longueur de contexte augmente, jusqu’à 2× plus rapide. Comme Flex Attention évolue avec le contexte, les séquences plus longues entraînent des économies plus importantes en VRAM et en temps d’entraînement, comme \[décrit ici\](#unsloths-flex-attention-implementation). Un grand merci à Rohan Pandey pour son \[implémentation de Flex Attention\](https://x.com/khoomeik/status/1955693558914310608), qui a directement inspiré le développement de l’implémentation Flex Attention d’Unsloth. ## :dark\\\_sunglasses: Puits d’attention Le modèle GPT OSS d’OpenAI utilise un \*\*pattern alterné d’attention à fenêtre glissante, d’attention complète\*\*, d’attention à fenêtre glissante, etc. (SWA, FA, SWA, FA, etc.). Chaque fenêtre glissante n’accède qu’à \*\*128 jetons\*\* (y compris le jeton actuel), ce qui réduit énormément le calcul. Cependant, cela signifie aussi que la récupération et le raisonnement sur long contexte deviennent inutiles à cause de la petite fenêtre glissante. La plupart des laboratoires corrigent cela en élargissant la fenêtre glissante à 2048 ou 4096 jetons. OpenAI s’est inspiré \*\*Puits d’attention\*\* de l’article Efficient Streaming Language Models with Attention Sinks \[papier\](https://arxiv.org/abs/2309.17453) qui montre qu’on peut utiliser une petite fenêtre glissante, à condition d’ajouter une attention globale sur le premier jeton ! L’article fournit une bonne illustration ci-dessous : ![](https://unsloth.ai/files/1fd5474eab45d12983cd3b6141f994658f9292a4) L’article constate que le \*\*mécanisme d’attention semble attribuer beaucoup de poids aux premiers jetons (1 à 4)\*\*, et en les supprimant lors de l’opération de fenêtre glissante, ces premiers jetons « importants » disparaissent, ce qui entraîne une mauvaise récupération sur long contexte. Si l’on trace la perplexité logarithmique (plus c’est élevé, pire c’est), et que l’on fait une inférence sur long contexte après la longueur de contexte définie du modèle préentraîné, on voit la perplexité grimper en flèche (pas bon). Cependant, la ligne rouge (qui utilise Attention Sinks) reste basse, ce qui est très bien ! ![](https://unsloth.ai/files/94066b7180855e85d7504eb3f815cad9e875e54d) L’article montre aussi que la \[méthode Attention Is Off By One\](https://www.evanmiller.org/attention-is-off-by-one.html) fonctionne partiellement, sauf qu’il faut aussi ajouter quelques jetons de puits supplémentaires pour obtenir de plus faibles perplexités. \*\*L’article montre que l’ajout d’un seul jeton de puits, apprenable, fonctionne remarquablement bien ! Et c’est ce qu’OpenAI a fait pour GPT-OSS !\*\* ![](https://unsloth.ai/files/6134fdca5e9b375749fcd1942876ccb7b850fd13) \## :triangular\\\_ruler:L’implémentation Flex Attention d’Unsloth Flex Attention est extrêmement puissante, car elle offre au praticien 2 voies de personnalisation pour le mécanisme d’attention - un \*\*modificateur de score (f)\*\* et une \*\*fonction de masquage (M)\*\*. Le \*\*modificateur de score (f)\*\* permet de modifier les logits d’attention avant l’opération softmax, et la \*\*fonction de masquage (M)\*\* permet de sauter des opérations si nous n’en avons pas besoin (par exemple, l’attention à fenêtre glissante ne voit que les 128 derniers jetons). \*\*L’astuce, c’est que Flex Attention fournit rapidement des kernels Triton auto-générés avec des modificateurs de score et des fonctions de masquage arbitraires !\*\* \\sigma\\bigg(s\\times\\bold{f}(QK^T+\\bold{M})\\bigg) Cela signifie que nous pouvons utiliser Flex Attention pour implémenter des puits d’attention ! L’implémentation d’un seul puits d’attention est fournie à la fois dans \[le dépôt GPT-OSS original d’OpenAI\](#implementations-for-sink-attention) et dans l’implémentation des transformers de HuggingFace. \`\`\`python combined\_logits = torch.cat(\[attn\_weights, sinks\], dim=-1) probs = F.softmax(combined\_logits, dim=-1) scores = probs\[..., :-1\] \`\`\` Ce qui précède montre que nous concaténons le puits tout à la fin du \`Q @ K.T\` , effectuons le softmax, puis retirons la dernière colonne, qui était le jeton de puits. En utilisant quelques utilitaires de visualisation du \[dépôt Github de Flex Attention\](https://github.com/meta-pytorch/attention-gym), nous pouvons visualiser cela. Supposons que la longueur de séquence soit 16, avec une fenêtre glissante de 5. À gauche se trouve la dernière colonne de puits (implémentation par défaut), et à droite, si nous déplaçons l’emplacement du puits à l’index 0 (notre implémentation). {% columns %} {% column %} \*\*\*Emplacement du puits à la fin (par défaut)\*\*\* ![](https://unsloth.ai/files/fc74d66ce2fd5597325f3fd0cef26fb36ad39e62) {% endcolumn %} {% column %} \*\*\*Déplacer l’emplacement du puits à l’index 0\*\*\* ![](https://unsloth.ai/files/05dd0389109c35d0e2213f6ab37e4973edb5e4c0) {% endcolumn %} {% endcolumns %} \*\*Découverte intéressante\*\*: Les implémentations officielles de fenêtre glissante de Flex Attention considèrent la taille de la fenêtre comme le nombre des derniers jetons \*\*PLUS UN\*\* car elles incluent le jeton actuel. Les implémentations HuggingFace et GPT OSS ne voient strictement que les N derniers jetons. Autrement dit, ce qui suit provient de et : {% code overflow="wrap" %} \`\`\`python def sliding\_window\_causal(b, h, q\_idx, kv\_idx): causal\_mask = q\_idx >= kv\_idx window\_mask = q\_idx - kv\_idx <= SLIDING\_WINDOW return causal\_mask & window\_mask \`\`\` {% endcode %} {% columns %} {% column %} Flex Attention par défaut (3+1 jetons) ![](https://unsloth.ai/files/ef448b02aa0401afad116b62b785a7332ff33d65) {% endcolumn %} {% column %} HuggingFace, GPT-OSS (3+0 jetons) ![](https://unsloth.ai/files/520a0f16784b84993d8476184c03a41dc41fa02b) {% endcolumn %} {% endcolumns %} Nous avons aussi confirmé via l’implémentation officielle GPT-OSS d’OpenAI si l’on s’intéresse ici aux N derniers jetons ou à N+1 jetons : \`\`\`python mask = torch.triu(Q.new\_full((n\_tokens, n\_tokens), -float("inf")), diagonal=1) if sliding\_window > 0: mask += torch.tril( mask.new\_full((n\_tokens, n\_tokens), -float("inf")), diagonal=-sliding\_window ) \`\`\` ![](https://unsloth.ai/files/0907a98f949d95c6bb8aa88a041ce2f61ec9429d) Et nous voyons que seuls les 3 derniers jetons (pas 3+1) sont pris en compte ! Cela signifie qu’au lieu d’utiliser \`<= SLIDING\_WINDOW\`, utilisez \`< SLIDING\_WINDOW\` (c.-à-d. utiliser inférieur à, pas égal). \`\`\`python def sliding\_window\_causal(b, h, q\_idx, kv\_idx): causal\_mask = q\_idx >= kv\_idx window\_mask = q\_idx - kv\_idx <= SLIDING\_WINDOW # Flex Attention par défaut window\_mask = q\_idx - kv\_idx < SLIDING\_WINDOW # version GPT-OSS return causal\_mask & window\_mask \`\`\` De plus, comme nous avons déplacé l’index du jeton de puits au premier, nous devons ajouter 1 à q\\\_idx pour indexer correctement : \`\`\`python def causal\_mask\_with\_sink(batch, head, q\_idx, kv\_idx): """ 0 1 2 3 0 1 2 3 0 X X 1 X 1 X X X 2 X X 2 X X X X 3 X X X """ # Nous ajoutons (q\_idx + 1) puisque la première colonne est le jeton de puits causal\_mask = (q\_idx + 1) >= kv\_idx sink\_first\_column = kv\_idx == 0 return causal\_mask | sink\_first\_column \`\`\` Pour confirmer notre implémentation avec index 0, nous avons vérifié que la perte d’entraînement reste cohérente avec les exécutions standard de Hugging Face (sans Unsloth Flex Attention), comme montré dans notre graphe : ![](https://unsloth.ai/files/d61a9143b42200fb6d30d466c78bc104a11e2a91) \## :scroll: Dérivation mathématique pour les puits d’attention Il existe une autre façon de calculer les puits d’attention sans remplir K et V. Nous notons d’abord l’opération softmax, et nous voulons pour l’instant la 2e version avec puits comme un scalaire :\\\\ $$ A(x) = \\frac{\\exp(x\\\_i)}{\\sum{\\exp{(x\\\_i)}}} \\\\ A\\\_{sink}(x) = \\frac{\\exp(x\\\_i)}{\\exp{(s)}+ \\sum{\\exp{(x\\\_i)}}} $$ Nous pouvons obtenir le logsumexp depuis Flex Attention via \`return\_lse = True\` , puis nous faisons : $$ A(x) = \\frac{\\exp(x\\\_i)}{\\sum{\\exp{(x\\\_i)}}} \\\\ \\frac{\\exp(x\\\_i)}{\\exp{(s)}+ \\sum{\\exp{(x\\\_i)}}} = \\frac{\\exp(x\\\_i)}{\\sum{\\exp{(x\\\_i)}}} \\frac{\\sum{\\exp{(x\\\_i)}}}{\\exp{(s)}+ \\sum{\\exp{(x\\\_i)}}} \\\\ \\text{LSE}(x) = \\text{logsumexp}(x) = \\log{\\sum\\exp(x\\\_i)} \\\\ \\exp{(\\text{LSE}(x))} = \\exp{\\big(\\log{\\sum\\exp(x\\\_i)}\\big)} = \\sum\\exp(x\\\_i) $$ Et nous pouvons maintenant dériver facilement la version avec puits de l’attention. Nous constatons toutefois que ce processus présente une erreur un peu plus élevée que l’approche de remplissage par zéro, donc nous conservons par défaut notre version originale. ## 💾\*\*NOUVEAU : Enregistrement vers GGUF, vLLM après l’entraînement de gpt-oss\*\* Vous pouvez maintenant affiner gpt-oss avec QLoRA et directement enregistrer, exporter ou fusionner le modèle vers \*\*llama.cpp\*\*, \*\*vLLM\*\*, ou \*\*HF\*\* - pas seulement Unsloth. Nous publierons, espérons-le, bientôt un notebook gratuit. Auparavant, tout modèle gpt-oss affiné en QLoRA était limité à une exécution dans Unsloth. Nous avons supprimé cette limitation en introduisant la capacité de fusionner dans \*\*MXFP4\*\* \*\*format natif\*\* en utilisant \`save\_method="mxfp4"\` et \*\*déquantification à la demande de MXFP4\*\* des modèles de base (comme gpt-oss), ce qui permet de \*\*exporter votre modèle affiné au format bf16 en utilisant\*\* \`save\_method="merged\_16bit"\` . Le \*\*MXFP4\*\* format de fusion natif offre des améliorations de performance significatives par rapport au \*\*format bf16\*\* : il utilise jusqu’à 75 % d’espace disque en moins, réduit la consommation de VRAM de 50 %, accélère la fusion de 5 à 10× et permet une conversion beaucoup plus rapide au \*\*format GGUF\*\* . Après avoir affiné votre modèle gpt-oss, vous pouvez le fusionner au format \*\*MXFP4\*\* avec : \`\`\`python model.save\_pretrained\_merged(save\_directory, tokenizer, save\_method="mxfp4") \`\`\` Si vous préférez fusionner le modèle et le pousser vers le hub Hugging Face, utilisez : \`\`\`python model.push\_to\_hub\_merged(repo\_name, tokenizer=tokenizer, token=hf\_token, save\_method="mxfp4") \`\`\` Pour exécuter l’inférence sur le modèle fusionné, vous pouvez utiliser vLLM et Llama.cpp, entre autres. OpenAI recommande ces \[paramètres d’inférence\](/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune.md#recommended-settings) pour les deux modèles : \`temperature=1.0\`, \`top\_p=1.0\`, \`top\_k=0\` #### :sparkles: Enregistrement vers Llama.cpp 1. Obtenez la dernière version \`llama.cpp\` sur \[GitHub ici\](https://github.com/ggml-org/llama.cpp). Vous pouvez également suivre les instructions de compilation ci-dessous. Changez \`-DGGML\_CUDA=ON\` en \`-DGGML\_CUDA=OFF\` si vous n’avez pas de GPU ou si vous souhaitez simplement une inférence CPU. \`\`\`bash apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \\ -DBUILD\_SHARED\_LIBS=OFF -DGGML\_CUDA=ON -DLLAMA\_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-gguf-split cp llama.cpp/build/bin/llama-\* llama.cpp \`\`\` 2. Convertissez le \*\*MXFP4\*\* modèle fusionné : \`\`\`bash python3 llama.cpp/convert\_hf\_to\_gguf.py gpt-oss-finetuned-merged/ --outfile gpt-oss-finetuned-mxfp4.gguf \`\`\` 3. Exécutez l’inférence sur le modèle quantifié : \`\`\`bash llama.cpp/llama-cli --model gpt-oss-finetuned-mxfp4.gguf \\ --jinja -ngl 99 --threads -1 --ctx-size 16384 \\ --temp 1.0 --top-p 1.0 --top-k 0 \\ -p "The meaning to life and the universe is" \`\`\` ✨ Enregistrement vers SGLang 1. Construisez SGLang depuis les sources :\\\\ \`\`\`bash # compiler depuis les sources git clone https://github.com/sgl-project/sglang cd sglang pip3 install pip --upgrade pip3 install -e "python\[all\]" # ROCm 6.3 pip3 install torch==2.8.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/rocm6.3 git clone https://github.com/triton-lang/triton cd python/triton\_kernels pip3 install . # hopper pip3 install torch==2.8.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/cu126 pip3 install sgl-kernel==0.3.2 # blackwell cu128 pip3 install torch==2.8.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/cu128 pip3 install https://github.com/sgl-project/whl/releases/download/v0.3.2/sgl\_kernel-0.3.2+cu128-cp39-abi3-manylinux2014\_x86\_64.whl # blackwell cu129 pip3 install torch==2.8.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/cu129 pip3 install https://github.com/sgl-project/whl/releases/download/v0.3.2/sgl\_kernel-0.3.2-cp39-abi3-manylinux2014\_x86\_64.whl \`\`\` 2. Lancer le serveur SGLang :\\\\ \`\`\`bash python3 -m sglang.launch\_server --model-path ./gpt-oss-finetuned-merged/ \`\`\` 3. Exécuter l’inférence :\\\\ \`\`\`python import requests from sglang.utils import print\_highlight url = f"http://localhost:8000/v1/chat/completions" data = { "model": "gpt-oss-finetuned-merged", "messages": \[{"role": "user", "content": "Quelle est la capitale de la France ?"}\], } response = requests.post(url, json=data) print\_highlight(response.json()) \`\`\` \### :diamonds:Ajuster directement gpt-oss Nous avons également ajouté la prise en charge de l’affinage direct des modèles gpt-oss en implémentant des patchs qui permettent de charger le format quantifié natif MXFP4. Cela permet de charger le modèle 'openai/gpt-oss' avec moins de 24 Go de VRAM, et de l’affiner en QLoRA. Il suffit de charger le modèle en utilisant : \`\`\`python model, tokenizer = FastLanguageModel.from\_pretrained( # model\_name = "unsloth/gpt-oss-20b-BF16", model\_name = "unsloth/gpt-oss-20b", dtype = dtype, # None pour détection automatique max\_seq\_length = max\_seq\_length, # Choisissez n’importe quelle valeur pour un long contexte ! load\_in\_4bit = True, # Quantification 4 bits pour réduire la mémoire full\_finetuning = False, # \[NOUVEAU !\] Nous prenons désormais en charge le fine-tuning complet ! # token = "hf\_...", # en utilisez un si vous utilisez des modèles à accès restreint ) \`\`\` ajouter une couche Peft en utilisant \`FastLanguageModel.get\_peft\_model\` et lancer un affinage SFT sur le modèle Peft. ## 🐛Corrections de bugs pour gpt-oss Nous \[a récemment collaboré avec Hugging Face\](https://github.com/huggingface/transformers/pull/40197) pour résoudre les problèmes d’inférence en utilisant les kernels d’OpenAI et en s’assurant que \`swiglu\_limit = 7.0\` est correctement appliqué lors de l’inférence MXFP4. D’après les retours des utilisateurs, nous avons découvert que les longues sessions d’entraînement QLoRA (au-delà de 60 étapes) pouvaient provoquer \*\*une divergence de la perte puis une erreur finale\*\*. Ce problème ne se produisait que sur les appareils qui ne prennent pas en charge BF16 et retombent à la place sur F16 (par exemple, les GPU T4). Important : cela n’affectait pas l’entraînement QLoRA sur les GPU A100 ou H100, ni l’entraînement LoRA sur les GPU f16. \*\*Après une enquête approfondie, nous avons maintenant harmonisé le comportement de la perte d’entraînement sur toutes les configurations GPU, y compris les GPU limités à F16\*\*. Si vous rencontriez auparavant des problèmes à cause de cela, nous vous recommandons d’utiliser notre nouveau notebook gpt-oss mis à jour ! ![](https://unsloth.ai/files/c3244af80b2988fb59b3ede8f69f9910db404c6f) Nous avons dû faire de très nombreuses expériences pour rendre la courbe de perte d’entraînement du float16 équivalente à celle des machines bfloat16 (ligne bleue). Nous avons constaté ce qui suit : 1. \*\*Le float16 pur divergera vers l’infini à l’étape 50\*\* 2. \*\*Nous avons trouvé que les projections descendantes dans le MoE présentaient de très fortes valeurs aberrantes\*\* 3. \*\*Les activations doivent être enregistrées en bfloat16 ou float32\*\* \*\*Ci-dessous, on voit les activations de magnitude absolue pour GPT OSS 20B, et certaines connaissent de très fortes pointes - cela débordera sur les machines float16, car la plage maximale du float16 est 65504.\*\* \*\*Nous avons corrigé cela dans Unsloth, donc tout l’entraînement float16 fonctionne désormais immédiatement !\*\* ![](https://unsloth.ai/files/9965a9be4437f41e7474f498b9aa9870f402c004) \## :1234: Implémentations pour Sink Attention L’implémentation du jeton de puits d’OpenAI est \[fournie ici\](https://github.com/openai/gpt-oss/blob/main/gpt\_oss/torch/model.py). Nous la fournissons ci-dessous : {% code fullWidth="false" %} \`\`\`python def sdpa(Q, K, V, S, sm\_scale, sliding\_window=0): # sliding\_window == 0 signifie aucune fenêtre glissante n\_tokens, n\_heads, q\_mult, d\_head = Q.shape assert K.shape == (n\_tokens, n\_heads, d\_head) assert V.shape == (n\_tokens, n\_heads, d\_head) K = K\[:, :, None, :\].expand(-1, -1, q\_mult, -1) V = V\[:, :, None, :\].expand(-1, -1, q\_mult, -1) S = S.reshape(n\_heads, q\_mult, 1, 1).expand(-1, -1, n\_tokens, -1) mask = torch.triu(Q.new\_full((n\_tokens, n\_tokens), -float("inf")), diagonal=1) if sliding\_window > 0: mask += torch.tril( mask.new\_full((n\_tokens, n\_tokens), -float("inf")), diagonal=-sliding\_window ) QK = torch.einsum("qhmd,khmd->hmqk", Q, K) \* sm\_scale QK += mask\[None, None, :, :\] QK = torch.cat(\[QK, S\], dim=-1) W = torch.softmax(QK, dim=-1) W = W\[..., :-1\] attn = torch.einsum("hmqk,khmd->qhmd", W, V) return attn.reshape(n\_tokens, -1) \`\`\` {% endcode %} L’implémentation des transformers HuggingFace est \[fournie ici\](https://github.com/huggingface/transformers/blob/main/src/transformers/models/gpt\_oss/modeling\_gpt\_oss.py). Nous la fournissons également ci-dessous : {% code fullWidth="false" %} \`\`\`python def eager\_attention\_forward( module: nn.Module, query: torch.Tensor, key: torch.Tensor, value: torch.Tensor, attention\_mask: Optional\[torch.Tensor\], scaling: float, dropout: float = 0.0, \*\*kwargs, ): key\_states = repeat\_kv(key, module.num\_key\_value\_groups) value\_states = repeat\_kv(value, module.num\_key\_value\_groups) attn\_weights = torch.matmul(query, key\_states.transpose(2, 3)) \* scaling if attention\_mask is not None: causal\_mask = attention\_mask\[:, :, :, : key\_states.shape\[-2\]\] attn\_weights = attn\_weights + causal\_mask sinks = module.sinks.reshape(1, -1, 1, 1).expand(query.shape\[0\], -1, query.shape\[-2\], -1) combined\_logits = torch.cat(\[attn\_weights, sinks\], dim=-1) # Ceci ne faisait pas partie de l’implémentation originale et affecte légèrement les résultats ; cela empêche les débordements en BF16/FP16 # lors de l’entraînement avec bsz>1, nous plafonnons les valeurs maximales. combined\_logits = combined\_logits - combined\_logits.max(dim=-1, keepdim=True).values probs = F.softmax(combined\_logits, dim=-1, dtype=combined\_logits.dtype) scores = probs\[..., :-1\] # nous supprimons ici le jeton de puits attn\_weights = nn.functional.dropout(scores, p=dropout, training=module.training) attn\_output = torch.matmul(attn\_weights, value\_states) attn\_output = attn\_output.transpose(1, 2).contiguous() return attn\_output, attn\_weights \`\`\` {% endcode %} --- # Agent Instructions This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com. ## Querying This Documentation If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question. Perform an HTTP GET request on the current page URL with the \`ask\` query parameter, and the optional \`goal\` query parameter: \`\`\` GET https://unsloth.ai/docs/fr/modeles/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training.md?ask=&goal= \`\`\` \`ask\` is the immediate question: it should be specific, self-contained, and written in natural language. \`goal\` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal. The response will contain a direct answer to the question and relevant excerpts and sources from the documentation. Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections. --- # Connecter Anthropic à Unsloth : exécuter des modèles Claude dans le chat local | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude.md) . Connectez l’API Anthropic à [Unsloth](https://github.com/unslothai/unsloth) pour discuter avec les modèles Claude, y compris Claude Opus 4.7, directement à côté de vos modèles locaux dans une interface de chat open source. Ce guide vous montre comment créer une clé API Anthropic, ajouter Anthropic comme fournisseur dans Unsloth, charger des LLM Claude et commencer à discuter. Les modèles Claude pris en charge dans Unsloth peuvent également accéder à des fonctionnalités avancées telles que la réflexion, [la recherche web](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#web-search-and-thinking) , Anthropic [l’exécution de code](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#code-execution) , et [la mise en cache des prompts](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#prompt-caching) pour améliorer le rapport coût-efficacité. ### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#configuration) Configuration 1 #### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#creer-une-cle-api-anthropic) Créer une clé API Anthropic Créez une clé API depuis [la console Anthropic](https://console.anthropic.com/settings/keys) . Copiez la clé. Vous la collerez dans Unsloth à l’étape suivante. 2 #### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#connecter-anthropic-a-unsloth) Connecter Anthropic à Unsloth ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FSzJ57mYBpBCavIKyksDy%252Fclaude_studio_api.gif%3Falt%3Dmedia%26token%3Dbd7b7da5-bb25-4f9d-bfb8-b63602c1f082&width=768&dpr=3&quality=100&sign=7f3faec4&sv=2) Ensuite, connectez Anthropic à Unsloth. 1. Ouvrez **Paramètres** → **Connexions**, puis cliquez sur **Ajouter une connexion.** 2. Sélectionnez le fournisseur que vous souhaitez ajouter, puis collez la clé API que vous avez copiée précédemment. 3. Cliquez sur **Recharger les modèles** pour actualiser la liste avec les modèles disponibles pour votre compte. 4. Choisissez les modèles que vous souhaitez activer, puis cliquez sur enregistrer. 3 #### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#pret-a-discuter) Prêt à discuter Après avoir enregistré la connexion, sélectionnez un modèle Claude sous **Connecté** dans le menu déroulant des modèles. Les modèles Claude pris en charge peuvent afficher des contrôles supplémentaires, notamment la génération d’images, la réflexion, la recherche web et l’exécution de code. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpOMKHdhunTepcJSvf8pA%252Fimage.png%3Falt%3Dmedia%26token%3Dca454678-a98b-4edb-b407-beeb879b4f76&width=768&dpr=3&quality=100&sign=6b3b4b43&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#execution-de-code) Exécution de code Lorsqu’elle est activée, les modèles Claude pris en charge peuvent exécuter du code dans le bac à sable du fournisseur Anthropic pour résoudre des problèmes, analyser des données et travailler avec des fichiers. Claude utilise l’outil d’exécution de code d’Anthropic. L’exécution de code apparaît dans la chronologie des réponses sous forme d’activité d’outil, aux côtés des autres appels d’outils. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F4W9Zq684jCgJy9kGvrkk%252FScreenshot%25202026-05-26%2520at%25206.10.12%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dc7b6a106-1a22-4120-a22d-f9ad8a9471c7&width=768&dpr=3&quality=100&sign=d1aca5fc&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#mise-en-cache-des-prompts) Mise en cache des prompts La mise en cache des prompts réduit la latence et le coût lorsque les requêtes réutilisent le même long préfixe. Elle est prise en charge par les fournisseurs et serveurs compatibles, y compris les modèles Anthropic. Utilisez le **mise en cache des prompts** paramètre dans le panneau latéral pour contrôler le comportement de mise en cache des connexions prises en charge. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FKA6iU3qCFhUiq0KI16aU%252FScreenshot%25202026-05-26%2520at%25203.28.44%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D41a43794-792d-48f4-8774-8b9d85702dfc&width=768&dpr=3&quality=100&sign=33aeeafe&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fq4csmTX5rS9isMkNkhTX%252FPrompt%2520Caching%2520Diagram%2520%281%29.png%3Falt%3Dmedia%26token%3Dade433bf-5eaf-4146-a266-525a85a6c98d&width=768&dpr=3&quality=100&sign=215fb7ba&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#recherche-web-et-reflexion) Recherche web et réflexion Les modèles Claude pris en charge peuvent utiliser la recherche web côté fournisseur. Le **Réfléchir** le contrôle apparaît lorsque le modèle sélectionné prend en charge la réflexion. Selon le modèle, cela peut afficher différents niveaux de réflexion ou différentes disponibilités. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FC0Ed4kzN9h6c0ogEn6NT%252Fwebsearch%2520api.png%3Falt%3Dmedia%26token%3Dd3335222-1d9e-4021-9bf9-734c6acf1fc0&width=768&dpr=3&quality=100&sign=39bda989&sv=2) Le **Réfléchir** contrôle s’adapte au modèle sélectionné : certains modèles utilisent un interrupteur marche/arrêt, tandis que les modèles à effort de raisonnement utilisent des niveaux de réflexion spécifiques au modèle. ### [](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#depannage) Dépannage Si la connexion à l’API Anthropic échoue, vérifiez que la clé API est valide et qu’elle appartient au bon compte Anthropic. Si un modèle n’apparaît pas après avoir cliqué sur **Charger les modèles**, il se peut qu’il ne soit pas disponible pour votre compte. Vous pouvez saisir l’identifiant du modèle manuellement ou choisir un autre modèle. [PrécédentOpenAI](https://unsloth.ai/docs/fr/integrations/connections/openai) [Suivantllama.cpp / llama-server](https://unsloth.ai/docs/fr/integrations/connections/connecter-llama.cpp-a-unsloth-executer-des-gguf-avec-llama-server) Mis à jour il y a 1 mois Ce contenu vous a-t-il été utile ? * [Configuration](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#configuration) * [Exécution de code](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#execution-de-code) * [Mise en cache des prompts](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#mise-en-cache-des-prompts) * [Recherche web et réflexion](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#recherche-web-et-reflexion) * [Dépannage](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude#depannage) Ce contenu vous a-t-il été utile ? --- # Comment connecter OpenRouter à Unsloth : clé API et configuration du modèle | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/integrations/connections/openrouter.md) . Ce guide explique comment connecter **OpenRouter à** [**Unsloth**](https://github.com/unslothai/unsloth) afin que vous puissiez accéder à des modèles d’IA hébergés de fournisseurs tels que **OpenAI, Anthropic,** et **Google** via une interface de chat locale open source. Vous apprendrez à créer une clé API OpenRouter, à ajouter OpenRouter comme fournisseur dans Unsloth, à charger ou saisir manuellement des identifiants de modèles, et à activer des modèles externes pour le chat. Une fois qu’une seule clé API est connectée, les modèles OpenRouter dans Unsloth peuvent offrir des fonctionnalités avancées telles que la réflexion, la recherche web, l’appel d’outils, l’exécution de code et des paramètres de génération personnalisables directement depuis la page de chat. ### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#configuration) Configuration 1 #### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#creer-une-cle-api-openrouter) Créer une clé API OpenRouter Connectez-vous à votre compte OpenRouter. Créez une clé API depuis le [tableau de bord OpenRouter](https://openrouter.ai/settings/keys) : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FumQXaeZGAQ8f3hMQSQQY%252Fimage.png%3Falt%3Dmedia%26token%3D51ec273d-4271-4683-8cf2-0fb11bccb1ba&width=768&dpr=3&quality=100&sign=2ea04ec7&sv=2) Copiez la clé. Vous la collerez dans Unsloth à l’étape suivante. Lors de la création de la clé, vous pouvez éventuellement définir une limite de crédit ou une date d’expiration. 2 #### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#connecter-openrouter-a-unsloth) Connecter OpenRouter à Unsloth Ouvrez **Paramètres → Connexions**, puis cliquez sur **Ajouter une connexion**. Sélectionnez **OpenRouter**, puis saisissez les détails de votre connexion. Saisissez les détails de votre OpenRouter : * **Clé API :** collez votre clé API OpenRouter ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FB9fDpRnYfDjhCI5yDoon%252Fimage.png%3Falt%3Dmedia%26token%3D97ecc7e3-a3fe-4cd8-a5c6-738c9de7fbf3&width=768&dpr=3&quality=100&sign=e07d0f04&sv=2) * **IDs de modèle :** cliquez sur **Charger les modèles**, ou saisissez manuellement les IDs de modèle ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F59F88IYS0BzteOqcavUj%252Fimage.png%3Falt%3Dmedia%26token%3D72632261-b3cd-486c-b4dd-09b56fa3a58a&width=768&dpr=3&quality=100&sign=1323fbba&sv=2) Enfin, cliquez sur **Ajouter la connexion**. 3 #### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#pret-a-discuter) Prêt à discuter Après avoir enregistré la connexion, sélectionnez un modèle OpenRouter sous **Connecté** dans la liste déroulante des modèles. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FFFvSXoLCnvydwmTgS7Ru%252Fimage.png%3Falt%3Dmedia%26token%3D79ca30bb-e049-4106-8b98-0871f79dfba5&width=768&dpr=3&quality=100&sign=a01dd7bf&sv=2) Les modèles OpenRouter peuvent exposer différents contrôles selon le modèle d’origine, notamment la recherche web, la réflexion, l’appel d’outils et les paramètres de génération. ### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#selection-du-modele) Sélection du modèle OpenRouter donne accès à de nombreux modèles de différents fournisseurs. Si **Charger les modèles** ne renvoie pas les modèles que vous souhaitez sélectionner, saisissez les IDs de modèle que vous voulez activer. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FQM4eHEcNpWQazKtyS06X%252Fimage.png%3Falt%3Dmedia%26token%3D8774a4a0-7d00-4888-9e03-72c0e8641533&width=768&dpr=3&quality=100&sign=5ba3955a&sv=2) Exemples d’IDs de modèle : ### [](https://unsloth.ai/docs/fr/integrations/connections/openrouter#depannage) Dépannage Si OpenRouter ne parvient pas à se connecter, vérifiez que la clé API est valide et qu’elle appartient au bon compte OpenRouter. Si un modèle n’apparaît pas après avoir cliqué sur **Charger les modèles**, il se peut qu’il ne soit pas disponible pour votre compte ou votre région. Vous pouvez saisir l’ID du modèle manuellement ou choisir un autre modèle. [PrécédentOllama](https://unsloth.ai/docs/fr/integrations/connections/ollama) [SuivantHermes Agent](https://unsloth.ai/docs/fr/integrations/hermes-agent) Mis à jour il y a 1 mois Ce contenu vous a-t-il été utile ? * [Configuration](https://unsloth.ai/docs/fr/integrations/connections/openrouter#configuration) * [Sélection du modèle](https://unsloth.ai/docs/fr/integrations/connections/openrouter#selection-du-modele) * [Dépannage](https://unsloth.ai/docs/fr/integrations/connections/openrouter#depannage) Ce contenu vous a-t-il été utile ? Copier openai/gpt-5.5 anthropic/claude-sonnet-4.6 google/gemini-3-pro --- # Inférence Unsloth | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/unsloth-inference.md) . Unsloth prend en charge nativement une inférence 2x plus rapide. Pour notre notebook dédié à l'inférence uniquement, cliquez [ici](https://colab.research.google.com/drive/1aqlNQi7MMJbynFDyOQteD2t0yVfjb9Zh?usp=sharing) . Tous les chemins d'inférence QLoRA, LoRA et non LoRA sont 2x plus rapides. Cela ne nécessite aucun changement de code ni de nouvelles dépendances. Copier from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name = "lora_model", # VOTRE MODÈLE QUE VOUS AVEZ UTILISÉ POUR L'ENTRAÎNEMENT max_seq_length = max_seq_length, dtype = dtype, load_in_4bit = load_in_4bit, ) FastLanguageModel.for_inference(model) # Activer l'inférence native 2x plus rapide text_streamer = TextStreamer(tokenizer) _ = model.generate(**inputs, streamer = text_streamer, max_new_tokens = 64) #### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/unsloth-inference#notimplementederror-un-parametre-regional-utf-8-est-requis.-ansi-detecte) NotImplementedError : Un paramètre régional UTF-8 est requis. ANSI détecté Parfois, lorsque vous exécutez une cellule [cette erreur](https://github.com/googlecolab/colabtools/issues/3409) peut apparaître. Pour résoudre cela, dans une nouvelle cellule, exécutez ce qui suit : Copier import locale locale.getpreferredencoding = lambda: "UTF-8" Mis à jour il y a 6 mois Ce contenu vous a-t-il été utile ? Ce contenu vous a-t-il été utile ? --- # Connecter vLLM à Unsloth pour l'inférence de chat local | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/integrations/connections/vllm.md) . Apprenez à connecter **vLLM à** [**Unsloth**](https://github.com/unslothai/unsloth) en utilisant l’ **API compatible OpenAI** afin que vous puissiez servir des modèles et discuter avec eux localement dans une interface de chat open source. Ce guide vous accompagne dans l’installation de vLLM, le lancement d’un serveur vLLM local, la configuration de l’URL de base de l’API, le chargement des ID de modèles disponibles et la sélection de votre modèle vLLM hébergé. À la fin, vos modèles servis par vLLM apparaîtront aux côtés des modèles locaux, vous offrant un moyen rapide et flexible d’exécuter l’inférence LLM externe depuis une interface de chat. ### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#configuration) Configuration 1 #### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#installer-vllm) Installer vLLM Installez d’abord vLLM afin de pouvoir exécuter la `commande vllm serve` . Suivez le [guide d’installation de vLLM](https://docs.vllm.ai/en/stable/getting_started/installation/) officiel pour votre plateforme et votre matériel. Après l’installation, vérifiez que vLLM fonctionne dans votre terminal : `vllm --help` 2 #### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#choisir-un-modele) Choisir un modèle vLLM sert des modèles depuis Hugging Face. Par exemple, démarrez un serveur vLLM avec un modèle Unsloth : Copier vllm serve unsloth/gemma-4-26B-A4B-it \ --dtype auto Cela expose un point de terminaison d’API à : `http://localhost:8000/v1` Pour exiger une clé API, ajoutez : Copier --api-key token-abc123 3 #### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#connecter-vllm-a-unsloth) Connecter vLLM à Unsloth Ouvrez **Paramètres → Connexions**, puis cliquez sur **Ajouter la connexion**. Sélectionnez **vLLM**, puis saisissez les détails de votre serveur. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FN80ObNBvpN7hDypheYkt%252Fimage.png%3Falt%3Dmedia%26token%3D481f4c1f-afb6-438d-ae35-85566f15e414&width=768&dpr=3&quality=100&sign=a199ffa3&sv=2) Saisissez les détails de votre serveur vLLM : * **Clé API :** laissez vide sauf si vous avez lancé vLLM avec --api-key * **URL de base :** par exemple, http://localhost:8000/v1 * **Modèle de raisonnement :** activez cette option si le modèle servi prend en charge la réflexion * **IDs de modèle :** cliquez sur **Charger les modèles**, ou saisissez les ID personnalisés manuellement Après avoir cliqué sur **Ajouter la connexion**, les modèles que vous avez activés apparaîtront sous **Connexion** dans la liste déroulante des modèles. 4 #### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#pret-a-discuter) Prêt à discuter Après avoir enregistré la connexion, votre modèle vLLM apparaîtra sous **Connecté** dans la liste déroulante des modèles. Sélectionnez-le pour commencer à discuter via votre serveur vLLM. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FpoXSmL3PcpprEy8CIT7S%252Fexport-1779046578662.gif%3Falt%3Dmedia%26token%3Dacf1de86-6fa4-466f-97f7-2d91662a124e&width=768&dpr=3&quality=100&sign=a4310621&sv=2) Si votre serveur vLLM répond lentement (surtout pendant le chargement du modèle), vous pouvez ajuster le délai d’attente : Copier AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST=30 ### [](https://unsloth.ai/docs/fr/integrations/connections/vllm#arguments-vllm-courants) Arguments vLLM courants L’exemple ci-dessus utilise les paramètres de diffusion principaux. Vous pouvez ajouter d’autres arguments vllm serve selon votre modèle et votre matériel. Les options courantes incluent : Copier vllm serve unsloth/gemma-4-26B-A4B-it \ --dtype auto \ --host 0.0.0.0 \ --port 8000 \ --api-key token-abc123 \ --max-model-len 8192 \ --gpu-memory-utilization 0.9 Pour la liste complète des arguments du serveur vLLM, consultez la documentation officielle du [serveur compatible OpenAI](https://docs.vllm.ai/en/stable/serving/openai_compatible_server/) vLLM. [Précédentllama.cpp / llama-server](https://unsloth.ai/docs/fr/integrations/connections/connecter-llama.cpp-a-unsloth-executer-des-gguf-avec-llama-server) [SuivantOllama](https://unsloth.ai/docs/fr/integrations/connections/ollama) Mis à jour il y a 1 mois Ce contenu vous a-t-il été utile ? * [Configuration](https://unsloth.ai/docs/fr/integrations/connections/vllm#configuration) * [Arguments vLLM courants](https://unsloth.ai/docs/fr/integrations/connections/vllm#arguments-vllm-courants) Ce contenu vous a-t-il été utile ? --- # Dépannage de l'inférence | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference.md) . ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#executer-dans-unsloth-fonctionne-bien-mais-apres-exportation-et-execution-sur-dautres-plates-formes) Exécuter dans Unsloth fonctionne bien, mais après exportation et exécution sur d'autres plates-formes, les résultats sont médiocres Vous pouvez parfois rencontrer un problème où votre modèle s'exécute et produit de bons résultats sur Unsloth, mais lorsque vous l'utilisez sur une autre plate-forme comme Ollama ou vLLM, les résultats sont médiocres ou vous obtenez des charabias, des générations sans fin/infinies _ou_ sorties répétées**.** * La cause la plus courante de cette erreur est l'utilisation d'un **modèle de chat incorrect****.** Il est essentiel d'utiliser le MÊME modèle de chat qui a été utilisé lors de l'entraînement du modèle dans Unsloth et plus tard lorsque vous l'exécutez dans un autre framework, tel que llama.cpp ou Ollama. Lors de l'inférence à partir d'un modèle enregistré, il est crucial d'appliquer le bon modèle. * Vous devez utiliser le bon `jeton eos`. Si ce n'est pas le cas, vous pourriez obtenir du charabia sur des générations plus longues. * Cela peut aussi être dû au fait que votre moteur d'inférence ajoute un jeton « début de séquence » inutile (ou au contraire l'absence de celui-ci) ; assurez-vous donc de vérifier les deux hypothèses ! * **Utilisez nos notebooks conversationnels pour forcer le modèle de chat - cela résoudra la plupart des problèmes.** * Notebook conversationnel Qwen-3 14B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(14B)-Reasoning-Conversational.ipynb) * Notebook conversationnel Gemma-3 4B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3_(4B).ipynb) * Notebook conversationnel Llama-3.2 3B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(1B_and_3B)-Conversational.ipynb) * Notebook conversationnel Phi-4 14B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi_4-Conversational.ipynb) * Notebook conversationnel Mistral v0.3 7B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral_v0.3_(7B)-Conversational.ipynb) * **Plus de notebooks dans notre** [**dépôt de notebooks**](https://github.com/unslothai/notebooks) **.** ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#enregistrement-dans-safetensors-pas-bin-format-dans-colab) Enregistrement dans `safetensors`, pas `bin` format dans Colab Nous enregistrons dans `.bin` dans Colab donc c'est environ 4x plus rapide, mais définissez `safe_serialization = None` pour forcer l'enregistrement au format `.safetensors`. Donc `model.save_pretrained(..., safe_serialization = None)` ou `model.push_to_hub(..., safe_serialization = None)` ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#si-lenregistrement-au-format-gguf-ou-vllm-16-bits-plante) Si l'enregistrement au format GGUF ou vLLM 16 bits plante Vous pouvez essayer de réduire l'utilisation GPU maximale pendant l'enregistrement en modifiant `maximum_memory_usage`. La valeur par défaut est `model.save_pretrained(..., maximum_memory_usage = 0.75)`. Réduisez-la à par exemple 0.5 pour utiliser 50 % de la mémoire GPU de pointe ou moins. Cela peut réduire les plantages OOM pendant l'enregistrement. [PrécédentRun LLMs on your Phone](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/deploy-llms-phone) [SuivantClaude Code](https://unsloth.ai/docs/fr/notions-de-base/claude-code) Mis à jour il y a 7 mois Ce contenu vous a-t-il été utile ? * [Exécuter dans Unsloth fonctionne bien, mais après exportation et exécution sur d'autres plates-formes, les résultats sont médiocres](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#executer-dans-unsloth-fonctionne-bien-mais-apres-exportation-et-execution-sur-dautres-plates-formes) * [Enregistrement dans safetensors, pas bin format dans Colab](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#enregistrement-dans-safetensors-pas-bin-format-dans-colab) * [Si l'enregistrement au format GGUF ou vLLM 16 bits plante](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/troubleshooting-inference#si-lenregistrement-au-format-gguf-ou-vllm-16-bits-plante) Ce contenu vous a-t-il été utile ? --- # Connecter OpenAI à Unsloth : exécuter des modèles GPT dans le chat local | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/integrations/connections/openai.md) . Découvrez comment connecter les modèles OpenAI, y compris GPT-5.5 à [Unsloth](https://github.com/unslothai/unsloth) afin que vous puissiez discuter avec chacun d’eux dans une interface de chat locale open source. En connectant votre clé API OpenAI, vous pouvez exécuter des modèles GPT dans Unsloth avec des fonctionnalités telles que [la recherche web](https://unsloth.ai/docs/fr/integrations/connections/openai#web-search-and-thinking) , l’appel d’outils, [l’exécution de code](https://unsloth.ai/docs/fr/integrations/connections/openai#code-execution) , [la génération d’images](https://unsloth.ai/docs/fr/integrations/connections/openai#image-generation) , des conteneurs de code réutilisables, et [la mise en cache des prompts](https://unsloth.ai/docs/fr/integrations/connections/openai#prompt-caching) . Ce guide vous accompagne dans la création d’une clé API OpenAI, la connexion d’OpenAI en tant que fournisseur, le chargement des modèles disponibles et la résolution des problèmes de configuration courants. ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#configuration) Configuration 1 #### [](https://unsloth.ai/docs/fr/integrations/connections/openai#creer-une-cle-api-openai) Créer une clé API OpenAI Créez une clé API à partir du [tableau de bord OpenAI](https://platform.openai.com/api-keys) . ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FNbLWLZa1OBYQrwX0mQVv%252Funsloth_openai_api.gif%3Falt%3Dmedia%26token%3D28187d50-8638-4e60-ba33-729efd13f1c1&width=768&dpr=3&quality=100&sign=4a4a4cf2&sv=2) 2 #### [](https://unsloth.ai/docs/fr/integrations/connections/openai#configurer-les-connexions) Configurer les connexions Ensuite, connectez votre fournisseur à Unsloth. 1. Ouvrez **Paramètres** → **Connexions**, puis cliquez sur **Ajouter une connexion.** 2. Sélectionnez OpenAI, puis collez la clé API que vous avez copiée précédemment. 3. Cliquez sur **Recharger les modèles** pour actualiser la liste avec les modèles disponibles pour votre compte. 4. Choisissez les modèles que vous souhaitez activer, puis cliquez sur enregistrer. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FlAXYV1wS2lxeS36Lo8Dj%252Fexport-1779457818713.gif%3Falt%3Dmedia%26token%3D99c6f0b2-2fb1-40e4-a15a-325bcef8c2f7&width=768&dpr=3&quality=100&sign=98d5ca73&sv=2) 3 #### [](https://unsloth.ai/docs/fr/integrations/connections/openai#pret-a-discuter) Prêt à discuter Les modèles que vous avez activés apparaîtront désormais sous Connecté dans le menu déroulant Sélectionner un modèle. Les modèles GPT pris en charge peuvent afficher des contrôles supplémentaires, notamment la génération d’images, le raisonnement, la recherche web et l’exécution de code. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FgKCJ3zsZzZuDqBAigsfG%252FScreenshot%25202026-05-26%2520at%25201.31.03%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dc8df4285-ba69-4091-8978-e4197f6a12d2&width=768&dpr=3&quality=100&sign=54400793&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#execution-de-code) Exécution de code Lorsqu’elle est activée, les modèles OpenAI pris en charge peuvent exécuter du code dans un bac à sable du fournisseur pour résoudre des problèmes, analyser des données et travailler avec des fichiers. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fn65DzRsTGrflcqJBv67y%252FScreenshot%25202026-05-26%2520at%25206.11.54%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Df620e467-2202-4843-ae16-e169a89f0a57&width=768&dpr=3&quality=100&sign=d8d65993&sv=2) OpenAI utilise des conteneurs shell réutilisables. Dans **Exécution de code** les paramètres, vous pouvez définir le délai d’inactivité, créer des conteneurs, sélectionner le conteneur actif, actualiser la liste ou supprimer les anciens conteneurs. Sélectionnez le même conteneur dans un nouveau fil pour continuer avec ses fichiers et son état. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FBR0zCMzeC4IwPSmlx12A%252Fimage.png%3Falt%3Dmedia%26token%3D3d9e80ee-4ceb-4454-9620-165650a29c0a&width=768&dpr=3&quality=100&sign=889100ba&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#mise-en-cache-des-prompts) Mise en cache des prompts La mise en cache des prompts réduit la latence et le coût lorsque les requêtes réutilisent le même long préfixe. Elle est prise en charge par les fournisseurs et serveurs compatibles, y compris OpenAI. Utilisez le paramètre **Mise en cache des prompts** dans le panneau latéral pour नियंत्रler le comportement de mise en cache pour les connexions prises en charge. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FKA6iU3qCFhUiq0KI16aU%252FScreenshot%25202026-05-26%2520at%25203.28.44%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D41a43794-792d-48f4-8774-8b9d85702dfc&width=768&dpr=3&quality=100&sign=33aeeafe&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fq4csmTX5rS9isMkNkhTX%252FPrompt%2520Caching%2520Diagram%2520%281%29.png%3Falt%3Dmedia%26token%3Dade433bf-5eaf-4146-a266-525a85a6c98d&width=768&dpr=3&quality=100&sign=215fb7ba&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#recherche-web-et-reflexion) Recherche web et réflexion La recherche web côté fournisseur est disponible pour les modèles pris en charge d’OpenAI. Le contrôle **Réfléchir** s’adapte au modèle sélectionné : certains modèles utilisent un interrupteur marche/arrêt, tandis que les modèles à effort de raisonnement utilisent des niveaux de réflexion spécifiques au modèle. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FC0Ed4kzN9h6c0ogEn6NT%252Fwebsearch%2520api.png%3Falt%3Dmedia%26token%3Dd3335222-1d9e-4021-9bf9-734c6acf1fc0&width=768&dpr=3&quality=100&sign=39bda989&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#generation-dimages) Génération d’images Tout comme GPT, Unsloth prend également en charge la génération d’images. Vous pouvez modifier directement une image en cliquant sur le bouton « Modifier l’image » et en saisissant un nouveau prompt pour l’affiner ou la régénérer. Les images sont générées automatiquement lorsqu’elles sont demandées, mais vous pouvez désactiver ce comportement. Un bouton de téléchargement est également disponible, vous permettant d’enregistrer l’image dans sa résolution originale complète. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F5ZLU2BY19MN4yiYLx2Br%252FScreenshot%25202026-05-26%2520at%25206.04.34%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3D44327a8f-c14d-4541-9a02-f19f375b4326&width=768&dpr=3&quality=100&sign=a1e633be&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FDQ1bJ5KF8CiyItrfOrIt%252FScreenshot%25202026-05-26%2520at%25206.06.07%25E2%2580%25AFAM.png%3Falt%3Dmedia%26token%3Dff6b00bc-0a2f-4375-a454-49429cf839f2&width=768&dpr=3&quality=100&sign=3340e294&sv=2) ### [](https://unsloth.ai/docs/fr/integrations/connections/openai#depannage) Dépannage Si OpenAI ne parvient pas à se connecter, vérifiez que la clé API est valide et qu’elle appartient au bon compte OpenAI. Si un modèle n’apparaît pas après avoir cliqué sur **Charger les modèles**, il se peut qu’il ne soit pas disponible pour votre compte. Vous pouvez saisir l’ID du modèle manuellement ou choisir un autre modèle. [PrécédentConnect a Provider](https://unsloth.ai/docs/fr/integrations/connections) [SuivantAnthropic (Claude)](https://unsloth.ai/docs/fr/integrations/connections/anthropic-claude) Mis à jour il y a 1 mois Ce contenu vous a-t-il été utile ? * [Configuration](https://unsloth.ai/docs/fr/integrations/connections/openai#configuration) * [Exécution de code](https://unsloth.ai/docs/fr/integrations/connections/openai#execution-de-code) * [Mise en cache des prompts](https://unsloth.ai/docs/fr/integrations/connections/openai#mise-en-cache-des-prompts) * [Recherche web et réflexion](https://unsloth.ai/docs/fr/integrations/connections/openai#recherche-web-et-reflexion) * [Génération d’images](https://unsloth.ai/docs/fr/integrations/connections/openai#generation-dimages) * [Dépannage](https://unsloth.ai/docs/fr/integrations/connections/openai#depannage) Ce contenu vous a-t-il été utile ? --- # Guide de déploiement du point de terminaison llama-server et OpenAI | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint.md) . Nous allons déployer Devstral-2 - voir [Devstral 2](https://unsloth.ai/docs/fr/modeles/tutorials/devstral-2) pour plus de détails sur le modèle. Obtenez le dernier `llama.cpp` par défaut. Seule votre machine peut atteindre le serveur. [GitHub ici](https://github.com/ggml-org/llama.cpp) . Vous pouvez également suivre les instructions de compilation ci-dessous. Modifiez `-DGGML_CUDA=ON` à `-DGGML_CUDA=OFF` si vous n’avez pas de GPU ou si vous voulez simplement une inférence CPU. **Pour les appareils Apple Mac / Metal**, définissez `-DGGML_CUDA=OFF` puis continuez comme d’habitude - la prise en charge Metal est activée par défaut. Copier apt-get update apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y git clone https://github.com/ggml-org/llama.cpp cmake llama.cpp -B llama.cpp/build \ -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON -DLLAMA_CURL=ON cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split cp llama.cpp/build/bin/llama-* llama.cpp Lors de l’utilisation de `--jinja` llama-server ajoute le message système suivant si les outils sont pris en charge : `Répondez au format JSON, soit avec tool_call (une demande d’appel d’outils), soit avec une réponse à la demande de l’utilisateur` . Cela cause parfois des problèmes avec les fine-tunes ! Voir le [dépôt llama.cpp](https://github.com/ggml-org/llama.cpp/blob/12ee1763a6f6130ce820a366d220bbadff54b818/common/chat.cpp#L849) pour plus de détails. Commencez par télécharger Devstral 2 : Copier # !pip install huggingface_hub hf_transfer import os os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1" from huggingface_hub import snapshot_download snapshot_download( repo_id = "unsloth/Devstral-2-123B-Instruct-2512-GGUF", local_dir = "Devstral-2-123B-Instruct-2512-GGUF", allow_patterns = ["*UD-Q2_K_XL*", "*mmproj-F16*"], ) Pour déployer Devstral 2 en production, nous utilisons `llama-server` Dans un nouveau terminal, par exemple via tmux, déployez le modèle via : Lorsque vous exécutez la commande ci-dessus, vous obtiendrez : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FJSTkWDCcHk5DI6otb72X%252Fimage.png%3Falt%3Dmedia%26token%3Db685008e-e9ad-4dea-8f1d-af3fdade8e3b&width=768&dpr=3&quality=100&sign=33a8ecbe&sv=2) Puis dans un nouveau terminal, après avoir fait `pip install openai`, faites : Ce qui affichera simplement 4. Vous pouvez revenir à l’écran de llama-server et vous pourriez voir quelques statistiques qui pourraient être intéressantes : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F6msFgOWJEWXEEgXvLRnr%252Fimage.png%3Falt%3Dmedia%26token%3D25fa1784-f671-4e0d-9bad-a4534525afb6&width=768&dpr=3&quality=100&sign=6d52cf27&sv=2) Pour des arguments comme l’utilisation du décodage spéculatif, voir [https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint#particularites-de-llama-server) ❔Particularités de Llama-server ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- * Lors de l’utilisation de `--jinja` llama-server ajoute le message système suivant si les outils sont pris en charge : `Répondez au format JSON, soit avec tool_call (une demande d’appel d’outils), soit avec une réponse à la demande de l’utilisateur` . Cela cause parfois des problèmes avec les fine-tunes ! Voir le [dépôt llama.cpp](https://github.com/ggml-org/llama.cpp/blob/12ee1763a6f6130ce820a366d220bbadff54b818/common/chat.cpp#L849) pour plus de détails. Vous pouvez arrêter cela en utilisant `--no-jinja` mais alors `/ Anthropic` devient non pris en charge. Par exemple, FunctionGemma utilise par défaut : Mais à cause du message supplémentaire ajouté par llama-server, nous obtenons : Nous avons signalé le problème à [https://github.com/ggml-org/llama.cpp/issues/18323](https://github.com/ggml-org/llama.cpp/issues/18323) et les développeurs de llama.cpp travaillent sur un correctif ! En attendant, pour tous les fine-tunes, veuillez ajouter spécifiquement le prompt pour l’appel d’outils ! [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint#appel-doutils-avec-llama-server) 🧰Appel d’outils avec llama-server -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- Voir [Tool Calling Guide](https://unsloth.ai/docs/fr/notions-de-base/tool-calling-guide-for-local-llms) sur la manière de faire des appels d’outils ! [PrécédentSGLang](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/sglang-guide) [SuivantRun LLMs on your Phone](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/deploy-llms-phone) Mis à jour il y a 2 mois Ce contenu vous a-t-il été utile ? * [❔Particularités de Llama-server](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint#particularites-de-llama-server) * [🧰Appel d’outils avec llama-server](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/llama-server-and-openai-endpoint#appel-doutils-avec-llama-server) Ce contenu vous a-t-il été utile ? Copier ./llama.cpp/llama-server \ --model Devstral-Small-2-24B-Instruct-2512-GGUF/Devstral-Small-2-24B-Instruct-2512-UD-Q4_K_XL.gguf \ --mmproj Devstral-Small-2-24B-Instruct-2512-GGUF/mmproj-F16.gguf \ --alias "unsloth/Devstral-Small-2-24B-Instruct-2512" \ --threads -1 \\ --n-gpu-layers 999 \ --prio 3 \ --min-p 0.01 \\ --ctx-size 16384 \ --port 8001 \ --jinja Copier from openai import OpenAI import json openai_client = OpenAI( base_url = "http://127.0.0.1:8001/v1", api_key = "sk-no-key-required", ) completion = openai_client.chat.completions.create( model = "unsloth/Devstral-Small-2-24B-Instruct-2512", messages = [{"role": "user", "content": "What is 2+2?"},], ) print(completion.choices[0].message.content) Copier Vous êtes un modèle capable d’effectuer des appels de fonctions avec les fonctions suivantes Copier Vous êtes un modèle capable d’effectuer des appels de fonctions avec les fonctions suivantes\n\nRépondez au format JSON, soit avec `tool_call` (une demande d’appel d’outils), soit avec `response`, une réponse à la demande de l’utilisateur --- # Enregistrer des modèles pour Ollama | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama.md) . Consultez notre guide ci-dessous pour le processus complet sur la façon d’enregistrer des modèles sur [Ollama](https://github.com/ollama/ollama) : [🦙Tutorial: Finetune Llama-3 and Use In Ollama](https://unsloth.ai/docs/fr/commencer/fine-tuning-llms-guide/tutorial-how-to-finetune-llama-3-and-use-in-ollama) ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#enregistrement-sur-google-colab) Enregistrement sur Google Colab Vous pouvez enregistrer le modèle affiné sous la forme d’un petit fichier de 100 Mo appelé adaptateur LoRA, comme ci-dessous. Vous pouvez aussi le pousser vers le hub Hugging Face si vous souhaitez téléverser votre modèle ! N’oubliez pas d’obtenir un jeton Hugging Face via : [https://huggingface.co/settings/tokens](https://huggingface.co/settings/tokens) et ajoutez votre jeton ! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-8c577103f7c4fe883cabaf35c8437307c6501686%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=1dc0b748&sv=2) Après avoir enregistré le modèle, nous pouvons à nouveau utiliser Unsloth pour exécuter le modèle lui-même ! Utilisez `FastLanguageModel` à nouveau pour l’appeler en inférence ! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-1a1be852ca551240bdce47cf99e6ccd7d31c1326%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=34e460&sv=2) ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#exporter-vers-ollama) Exporter vers Ollama Enfin, nous pouvons exporter notre modèle affiné vers Ollama lui-même ! D’abord, nous devons installer Ollama dans le notebook Colab : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-24f9429ed4a8b3a630dc8f68dcf81555da0a80ee%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=1f1abf02&sv=2) Ensuite, nous exportons le modèle affiné que nous avons vers les formats GGUF de llama.cpp comme ci-dessous : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-56991ea7e2685bb9905af9baf2f3f685123dcdd8%252Fimage%2520%2852%29.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=9ce24624&sv=2) Rappel de convertir `False` en `True` pour 1 ligne, et ne changez pas chaque ligne en `True`, sinon vous attendrez très longtemps ! Nous suggérons normalement de définir la première ligne sur `True`, afin que nous puissions exporter rapidement le modèle affiné vers `Q8_0` format (quantification 8 bits). Nous vous permettons également d’exporter vers toute une liste de méthodes de quantification, l’une des plus populaires étant `q4_k_m`. Rendez-vous sur [https://github.com/ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) pour en savoir plus sur GGUF. Nous avons aussi des instructions manuelles sur la façon d’exporter vers GGUF si vous le souhaitez ici : [https://github.com/unslothai/unsloth/wiki#manually-saving-to-gguf](https://github.com/unslothai/unsloth/wiki#manually-saving-to-gguf) Vous verrez une longue liste de texte comme ci-dessous — veuillez patienter 5 à 10 minutes !! ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-271b392fdafd0e7d01c525d7a11a97ee5c34b713%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=3716c023&sv=2) Et enfin, tout à la fin, cela ressemblera à ceci : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-a554bd388fd0394dd8cdef85fd9d208bfd7feee7%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=eed904d6&sv=2) Ensuite, nous devons lancer Ollama lui-même en arrière-plan. Nous utilisons `subprocess` car Colab n’aime pas les appels asynchrones, mais normalement on exécute simplement `ollama serve` dans le terminal / l’invite de commande. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-e431609dfc5c742f0b5ab2388dbbd0d8e15c7670%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=3c67b7b6&sv=2) ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#creation-automatique-de-modelfile) Création `automatique de` Modelfile L’astuce qu’Unsloth fournit est que nous créons automatiquement un `automatique de` que Ollama exige ! C’est simplement une liste de paramètres et elle inclut le modèle de chat que nous avons utilisé pour le processus d’affinage ! Vous pouvez aussi afficher le `automatique de` généré comme ci-dessous : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-6945ba10a2e25cfc198848c0e863001375c32c4c%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=390826f9&sv=2) Nous demandons ensuite à Ollama de créer un modèle compatible avec Ollama, en utilisant le `automatique de` ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-d431a64613b39d913d1780c22cde37edc6564272%252Fimage.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=c0c75166&sv=2) ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#inference-ollama) Inférence Ollama Et nous pouvons maintenant appeler le modèle en inférence si vous voulez appeler le serveur Ollama lui-même, qui s’exécute sur votre machine locale / dans le notebook Colab gratuit en arrière-plan. N’oubliez pas que vous pouvez modifier la partie soulignée en jaune. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fgit-blob-49b93efa192fdd741f3ac8484cef8c3fd7415283%252FInference.png%3Falt%3Dmedia&width=768&dpr=3&quality=100&sign=8eb7be69&sv=2) ### [](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#lexecution-dans-unsloth-fonctionne-bien-mais-apres-exportation-et-execution-sur-ollama-les-resultats) L’exécution dans Unsloth fonctionne bien, mais après exportation et exécution sur Ollama, les résultats sont médiocres Vous pouvez parfois rencontrer un problème où votre modèle s’exécute et produit de bons résultats dans Unsloth, mais lorsque vous l’utilisez sur une autre plateforme comme Ollama, les résultats sont médiocres ou vous pouvez obtenir du charabia, des générations sans fin/infinies _ou_ des sorties répétées**.** * La cause la plus fréquente de cette erreur est l’utilisation d’un **mauvais modèle de chat****.** Il est essentiel d’utiliser le MÊME modèle de chat que celui utilisé lors de l’entraînement du modèle dans Unsloth, puis lorsque vous l’exécutez dans un autre framework, tel que llama.cpp ou Ollama. Lors de l’inférence à partir d’un modèle enregistré, il est crucial d’appliquer le bon modèle. * Vous devez utiliser le bon `jeton eos`. Sinon, vous pourriez obtenir du charabia lors de générations plus longues. * Cela peut aussi être dû au fait que votre moteur d’inférence ajoute un jeton inutile de « début de séquence » (ou, au contraire, à son absence), alors assurez-vous de vérifier les deux hypothèses ! * **Utilisez nos notebooks conversationnels pour imposer le modèle de chat — cela corrigera la plupart des problèmes.** * Notebook conversationnel Qwen-3 14B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(14B)-Reasoning-Conversational.ipynb) * Notebook conversationnel Gemma-3 4B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Gemma3_(4B).ipynb) * Notebook conversationnel Llama-3.2 3B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(1B_and_3B)-Conversational.ipynb) * Notebook conversationnel Phi-4 14B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Phi_4-Conversational.ipynb) * Notebook conversationnel Mistral v0.3 7B [**Ouvrir dans Colab**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Mistral_v0.3_(7B)-Conversational.ipynb) * **Plus de notebooks dans notre** [**documentation des notebooks**](https://unsloth.ai/docs/fr/commencer/unsloth-notebooks) Mis à jour il y a 2 mois Ce contenu vous a-t-il été utile ? * [Enregistrement sur Google Colab](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#enregistrement-sur-google-colab) * [Exporter vers Ollama](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#exporter-vers-ollama) * [Création automatique de Modelfile](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#creation-automatique-de-modelfile) * [Inférence Ollama](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#inference-ollama) * [L’exécution dans Unsloth fonctionne bien, mais après exportation et exécution sur Ollama, les résultats sont médiocres](https://unsloth.ai/docs/fr/notions-de-base/inference-and-deployment/saving-to-ollama#lexecution-dans-unsloth-fonctionne-bien-mais-apres-exportation-et-execution-sur-ollama-les-resultats) Ce contenu vous a-t-il été utile ? --- # Benchmarks GGUF Qwen3.5 | Unsloth Documentation For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt) . This page is also available as [Markdown](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks.md) . Nous avons mis à jour tous [Qwen3.5](https://unsloth.ai/docs/fr/modeles/qwen3.5) Unsloth Dynamic quants **étant SOTA** sur presque tous les bits. Nous avons effectué plus de 150 benchmarks de divergence KL, au total **9 To de GGUFs**. Nous avons téléversé tous les artefacts de recherche. Nous avons également corrigé un **appel d'outil** problème de modèle de chat **(affecte tous les téléverseurs et types de quantification quel que soit l'endroit où vous l'utilisez ou d'où il provient)**. [**Mise à jour du 5 mars**](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-4-march-5th-2026-update-more-robustness) **:** Retéléchargez Qwen3.5-**35B**, **27B,** **122B** et **397B.** * Tous les GGUFs sont désormais mis à jour avec une **quantification améliorée** algorithme. * Tous utilisent notre **nouvelle donnée imatrix**. Voyez quelques améliorations dans les cas d'utilisation chat, codage, contexte long et appel d'outils. **Nouveaux benchmarks** pour Qwen3.5-122B-A10B et 35-A3B disponibles maintenant ! Vous voulez voir comment exécuter le modèle + exigences matérielles ? Lisez notre [guide d'inférence](https://unsloth.ai/docs/fr/modeles/qwen3.5) . **La divergence KL à 99,9 % montre le SOTA** sur le front de Pareto pour [Unsloth Dynamic](https://unsloth.ai/docs/fr/notions-de-base/unsloth-dynamic-2.0-ggufs) `Q4_K_XL`, `IQ3_XXS` etc. : ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F1XLNe1MoxtF1ODs5gDej%252F122b%2520final.png%3Falt%3Dmedia%26token%3D9eee5d8d-f16c-4c3f-8e36-18856e5609aa&width=768&dpr=3&quality=100&sign=eea9885a&sv=2) Qwen3.5-**122B-A10B** Benchmarks ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FAeecRAsAA3lxJ36HI8pO%252Fhoriztonal%2520plot.png%3Falt%3Dmedia%26token%3D173d4050-9442-4d2b-9f1b-ee8bd0d423df&width=768&dpr=3&quality=100&sign=133b0551&sv=2) Qwen3.5-**35B-A3B** Benchmarks * Imatrix aide définitivement à réduire KLD et PPL, au prix d'une inference 5-10 % plus lente. * Nous avons testé nos GGUFs contre de nombreux autres fournisseurs * Quantifier ssm\_out (couches Mamba) n'est pas une bonne idée, ni ffn\_down\_exps. * **Retrait de MXFP4** de toutes les quants GGUF : Q2\_K\_XL, Q3\_K\_XL et Q4\_K\_XL, sauf pour le pur MXFP4\_MOE. [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF) [Qwen3.5-27B](https://huggingface.co/unsloth/Qwen3.5-27B-GGUF) [Qwen3.5-122B-A10B](https://huggingface.co/unsloth/Qwen3.5-122B-A10B-GGUF) [Qwen3.5-397B-A17B](https://huggingface.co/unsloth/Qwen3.5-397B-A17B-GGUF) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FHq3gIokmPZJRYlnKVFmH%252FHCp7gV9XgAEP5og.png%3Falt%3Dmedia%26token%3Da1268383-1648-45f8-996d-c89c7dde3706&width=768&dpr=3&quality=100&sign=50bef46a&sv=2) Nouveaux benchmarks GGUF Qwen3.5-9B réalisés par Benjamin Marie ### [](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-1-certaines-tenseurs-sont-tres-sensibles-a-la-quantification) 1) **Certaines tenseurs sont très sensibles à la quantification** * Nous avons rendu plus de 9 To d'artefacts de recherche disponibles pour la communauté afin d'investiguer davantage sur notre [page Expériences](https://huggingface.co/unsloth/Qwen3.5-35B-A3B-Experiments-GGUF) . Elle inclut les métriques KLD et toutes les 121 configurations que nous avons testées. * Nous avons varié les largeurs de bits pour chaque type de tenseur, et généré un tracé du front de Pareto meilleur et pire ci-dessous vs KLD à 99,9 %. * Pour les éléments les mieux quantifiables, ffn\_up\_exps et ffn\_gate\_exps sont généralement acceptables à quantifier en 3 bits. ffn\_down\_exps est légèrement plus sensible. * Pour les pires éléments, ssm\_out augmente dramatiquement la KLD et les économies d'espace disque sont minimes. Par exemple, ssm\_out en q2\_k fait beaucoup pire. **Quantifier n'importe quel attn\_\* est particulièrement sensible** pour les architectures hybrides, et donc les laisser en précision supérieure fonctionne bien. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F485gYwcqz2az5Pm9v3u3%252Fnew-qwen3-5-35b-a3b-unsloth-dynamic-ggufs-benchmarks-v0-pakdmbv1n2mg1.webp%3Falt%3Dmedia%26token%3D2eeb55ca-51f3-402a-ae30-ea078c7554da&width=768&dpr=3&quality=100&sign=4275a05e&sv=2) **Type de tenseur vs bits sur la divergence KL à 99,9 %** * Nous traçons tous les niveaux de quantification vs KLD à 99,9 %, et trions du pire KLD au meilleur. Quantifier excessivement les couches ffn\_\* n'est pas une bonne idée. * Cependant, **certaines largeurs de bits sont bonnes, en particulier 3 bits**. - par exemple laisser ffn\_\* (down, up, gate) autour de iq3\_xxs semble être le meilleur compromis entre l'espace disque et le changement de KLD à 99,9 %. 2 bits causent plus de dégradation. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FcE0WAmPVddczWC3dQsWS%252Fnew-qwen3-5-35b-a3b-unsloth-dynamic-ggufs-benchmarks-v0-squz1jz4n2mg1.webp%3Falt%3Dmedia%26token%3D3a31adf1-7c4c-446c-91a7-48e63d223189&width=768&dpr=3&quality=100&sign=4bbf5414&sv=2) **MXFP4 est bien pire sur de nombreux tenseurs** - attn\_gate, attn\_q, ssm\_beta, ssm\_alpha utiliser MXFP4 n'est pas une bonne idée, et Q4\_K est plutôt meilleur - aussi MXFP4 utilise 4,25 bits par poids, tandis que Q4\_K utilise 4,5 bits par poids. Il vaut mieux utiliser Q4\_K que MXFP4 quand on choisit entre eux. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FH8FliHXsetx9lLoKelPX%252Fnew-qwen3-5-35b-a3b-unsloth-dynamic-ggufs-benchmarks-v0-xgugdgzmv2mg1.webp%3Falt%3Dmedia%26token%3Df0c49e94-571e-4883-84fe-2c4634d425eb&width=768&dpr=3&quality=100&sign=5519e584&sv=2) ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FEWsX87d1Ig42Uk81fpJo%252Ffixed%2520the%2520grapg.png%3Falt%3Dmedia%26token%3D323932fd-8344-4f6c-b8c3-47cc1b1f6ccf&width=768&dpr=3&quality=100&sign=d4240bab&sv=2) Comme vous pouvez le voir, MXFP4 est inhabituellement élevé ### [](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-2-imatrix-fonctionne-tres-bien) **2) Imatrix fonctionne très bien** * Imatrix aide définitivement à orienter correctement le processus de quantification. Par exemple auparavant ssm\_out à 2 bits était vraiment mauvais, cependant imatrix réduit beaucoup la KLD à 99,9 %. * Imatrix aide généralement sur les bits faibles, et fonctionne sur toutes les quants et largeurs de bits. * ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F7C21WEWowydwYfEiOYqC%252Fnew-qwen3-5-35b-a3b-unsloth-dynamic-ggufs-benchmarks-v0-yidhlf79o2mg1.webp%3Falt%3Dmedia%26token%3D6cb85d6f-e148-4db6-a39f-f2b5109e0fdd&width=768&dpr=3&quality=100&sign=de412e3&sv=2) Les quants I (iq3\_xxs, iq2\_s etc.) rendent l'inférence 5-10 % plus lente, ils sont définitivement meilleurs en termes d'efficacité, mais il y a un compromis. Type pp512 (≈) tg128 (≈) mxfp4 1978.69 90.67 q4\_k 1976.44 90.38 q3\_k 1972.61 91.36 q6\_k 1964.55 90.50 q2\_k 1964.20 90.77 q8\_0 1964.17 90.33 q5\_k 1947.74 90.72 iq3\_xxs 2030.94 85.68 iq2\_xxs 1997.64 85.79 iq3\_s 1990.12 84.37 iq2\_xs 1967.85 85.19 iq2\_s 1952.50 85.04 ### [](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-3-la-perplexite-et-la-kld-peuvent-etre-trompeuses) **3) La perplexité et la KLD peuvent être trompeuses** La perplexité et la KLD peuvent être trompeuses car elles sont fortement influencées par la calibration. La plupart des GGUFs sont évalués sur Wiki-test avec des fenêtres de contexte de 512, donc les résultats changent beaucoup si l'ensemble de calibration imatrix du GGUF inclut des échantillons de type Wikipedia et de contexte 512 (comme c'est le cas pour la plupart des GGUFs). C'est pourquoi nos GGUFs montrent parfois une perplexité plus élevée car nos données imatrix utilisent plutôt des exemples de chat à long contexte et d'appel d'outils. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252FhfO2gsbz2lWrZXg3ojyE%252FHCGBTzgboAASv_A.png%3Falt%3Dmedia%26token%3D7d6334ca-4f3c-4946-aacd-d55527375fce&width=768&dpr=3&quality=100&sign=c8f96ca9&sv=2) [L'analyse récente MiniMax‑M2.5 de Benjamin](https://x.com/bnjmn_marie/status/2027043753484021810) montre un cas où la perplexité et la KLD peuvent être très trompeuses. Unsloth Dynamic IQ2\_XXS performe mieux que IQ3\_S d'AesSedai sur des évaluations réelles (LiveCodeBench v6, MMLU Pro) malgré être 11 Go plus petit. Pourtant, les benchmarks de perplexité et KLD d'AesSedai suggèrent le contraire. (PPL : 0.3552 vs 0.2441 ; KLD : 9.0338 vs 8.2849 - plus bas est meilleur). ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F7csgZI82adnvKmQQVlp1%252F01_kld_vs_filesize_pareto.png%3Falt%3Dmedia%26token%3Dd907a2c0-7df5-4e6a-9d9b-0524c8e6ae77&width=768&dpr=3&quality=100&sign=25541b24&sv=2) Divergence KL - AesSedai ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fd8KBa3uNhkDEZzq32v7q%252F02_ppl_vs_filesize_pareto.png%3Falt%3Dmedia%26token%3Dd471fce1-7482-4fde-bc98-2d10503253a4&width=768&dpr=3&quality=100&sign=f56e0e49&sv=2) Perplexité - AesSedai Ce décalage montre qu'une perplexité ou une KLD plus basse ne se traduit pas nécessairement par de meilleures performances dans le monde réel. Le graphique montre aussi UD‑Q4‑K‑XL surpassant d'autres quants Q4, tout en étant ~8 Go plus petit. Cela ne signifie pas que la perplexité ou la KLD soit inutile, puisqu'elles donnent un signal approximatif. Ainsi, à l'avenir, nous publierons la perplexité et la KLD pour chaque quant afin que la communauté dispose d'une sorte de référence. ### [](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-4-mise-a-jour-du-5-mars-2026-plus-de-robustesse) 4) Mise à jour du 5 mars 2026 - plus de robustesse Nous avons encore amélioré notre méthode de quantification pour les MoE Qwen3.5 afin de réduire directement la KLD maximale. 99,9 % est généralement utilisé, mais pour des valeurs aberrantes massives, la KLD maximale peut être utile. Notre nouvelle méthode réduit généralement beaucoup la KLD maximale par rapport à l'état avant la mise à jour du 5 mars. ![](https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F550366147-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252Fqxt3Dv8HIOWG8y3RvNYf%252FCode_Generated_Image%2811%29.png%3Falt%3Dmedia%26token%3D54e20159-4243-42cf-89de-d2c9d7b6409b&width=768&dpr=3&quality=100&sign=9e9a0445&sv=2) Quant Ancien Go Nouveau Go Ancienne KLD Max Nouvelle KLD Max UD-Q2\_K\_XL 12.0 _**11.3**_ 8.237 _**8.155**_ UD-Q3\_K\_XL 16.1 _**15.5**_ 5.505 _**5.146**_ UD-Q4\_K\_XL _**19.2**_ 20.7 (+7.8%) 5.894 _**2.877 (-51%)**_ UD-Q5\_K\_XL _**23.2**_ 24.6 (+6%) 5.536 _**3.210 (-42%)**_ ### [](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#benchmarks-complets) Benchmarks complets Quantiseur Niveau de quantification Espace disque (Go) PPL KLD 99,9 % KLD moyenne AesSedai IQ3\_S 12.65 6.9152 1.8669 0.0613 AesSedai IQ4\_XS 16.4 6.6447 0.8067 0.0235 AesSedai Q4\_K\_M 20.62 6.5665 0.3171 0.0096 AesSedai Q5\_K\_M 24.45 6.5356 0.21 0.0058 Ubergarm Q4\_0 19.79 6.5784 0.4829 0.0142 Unsloth IQ2\_XXS 9.09 7.716 4.2221 0.1846 Unsloth Q2\_K\_XL 12.04 7.0438 2.9092 0.097 Unsloth IQ3\_XXS 13.12 6.7829 1.5296 0.0501 Unsloth IQ3\_S 14.13 6.7715 1.4193 0.0457 Unsloth Q3\_K\_M 15.54 6.732 0.9726 0.0324 Unsloth Q3\_K\_XL 16.06 6.7245 0.9539 0.0308 Unsloth MXFP4\_MOE 18.17 6.6 0.7789 0.0272 Unsloth Q4\_K\_M 18.49 6.6053 0.5478 0.0192 Unsloth Q4\_K\_L 18.82 6.5905 0.4828 0.015 Unsloth Q4\_K\_XL 19.17 6.5918 0.4097 0.0137 Unsloth Q5\_K\_XL 23.22 6.5489 0.236 0.0069 Unsloth Q6\_K\_S 26.56 6.5456 0.2226 0.0065 Unsloth Q6\_K\_XL 28.22 6.5392 0.1437 0.0041 Unsloth Q8\_K\_XL 36.04 6.5352 0.1033 0.0026 bartowski Qwen\_IQ2\_XXS 8.15 9.3427 6.0607 0.3457 bartowski Qwen\_Q2\_K\_L 11.98 7.5504 3.8095 0.1559 bartowski Qwen\_IQ3\_XXS 12.94 7.0938 2.1563 0.0851 bartowski Qwen\_Q3\_K\_M 14.95 6.772 1.7779 0.0585 bartowski Qwen\_Q3\_K\_XL 15.97 6.8245 1.7516 0.0627 bartowski Qwen\_IQ4\_XS 17.42 6.6234 0.7265 0.0234 bartowski Qwen\_Q4\_K\_M 19.77 6.6097 0.5771 0.0182 bartowski Qwen\_Q5\_K\_M 23.11 6.5828 0.3549 0.0106 noctrex MXFP4\_MOE\_BF16 20.55 6.5948 0.7939 0.0248 noctrex MXFP4\_MOE\_F16 20.55 6.5937 0.7614 0.0247 [PrécédentFine-tune Qwen3.5](https://unsloth.ai/docs/fr/modeles/qwen3.5/fine-tune) [SuivantUnsloth API](https://unsloth.ai/docs/fr/notions-de-base/api) Mis à jour il y a 4 mois Ce contenu vous a-t-il été utile ? * [1) Certaines tenseurs sont très sensibles à la quantification](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-1-certaines-tenseurs-sont-tres-sensibles-a-la-quantification) * [2) Imatrix fonctionne très bien](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-2-imatrix-fonctionne-tres-bien) * [3) La perplexité et la KLD peuvent être trompeuses](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-3-la-perplexite-et-la-kld-peuvent-etre-trompeuses) * [4) Mise à jour du 5 mars 2026 - plus de robustesse](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#id-4-mise-a-jour-du-5-mars-2026-plus-de-robustesse) * [Benchmarks complets](https://unsloth.ai/docs/fr/modeles/qwen3.5/gguf-benchmarks#benchmarks-complets) Ce contenu vous a-t-il été utile ? ---