If you read our graphic, it says you can update the template as well. Most people don't know how to replace the chat template.
Daniel (Unsloth) PRO
AI & ML interests
None yet
Recent Activity
repliedto their post about 3 hours ago
Gemma 4 is now faster and much more accurate! ๐
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4 liked a model about 7 hours ago
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF updated a Space about 22 hours ago
unsloth/READMEOrganizations
replied to their post about 3 hours ago
replied to their post 1 day ago
Our MLX quants were update: https://huggingface.co/collections/unsloth/gemma-4
replied to their post 1 day ago
It was posted officially by Google: https://x.com/googlegemma/status/2077449152062247219
posted an update 2 days ago
Post
4382
Gemma 4 is now faster and much more accurate! ๐
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
posted an update 6 days ago
Post
4366
Weโre releasing Gemma 4 NVFP4 quants that run 1.5ร faster on your GPU.
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
posted an update 10 days ago
Post
4042
Weโre releasing new Qwen3.6 quants that run 2.5ร faster on your GPU. โก
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4
posted an update 13 days ago
Post
5960
DeepSeek-V4 can now run locally with Unsloth GGUFs! ๐ณ
Run lossless DeepSeek-V4-Flash on 168GB RAM or
3-bit works on 110GB Mac, RAM, VRAM setups.
Run via Unsloth Studio or llama.cpp.
GGUF: unsloth/DeepSeek-V4-Flash-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
Run lossless DeepSeek-V4-Flash on 168GB RAM or
3-bit works on 110GB Mac, RAM, VRAM setups.
Run via Unsloth Studio or llama.cpp.
GGUF: unsloth/DeepSeek-V4-Flash-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
posted an update 27 days ago
Post
3317
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5
We gave 3 models the same prompt and compared one-shot outputs.
The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.
Which output do you like best?
GGUF: unsloth/GLM-5.2-GGUF
We gave 3 models the same prompt and compared one-shot outputs.
The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.
Which output do you like best?
GGUF: unsloth/GLM-5.2-GGUF
posted an update about 1 month ago
Post
4589
Google's new DiffusionGemma can now run at 2000+ tokens/sec! โก
We made local DiffusionGemma inference 1.8ร faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/models/diffusiongemma
We made local DiffusionGemma inference 1.8ร faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/models/diffusiongemma
posted an update about 1 month ago
Post
1182
Google releases DiffusionGemma.โจ
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
Run with 4x faster text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio.
GGUF: unsloth/diffusiongemma-26B-A4B-it-GGUF
Guide: https://unsloth.ai/docs/models/diffusiongemma
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
Run with 4x faster text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio.
GGUF: unsloth/diffusiongemma-26B-A4B-it-GGUF
Guide: https://unsloth.ai/docs/models/diffusiongemma
posted an update about 1 month ago
Post
4289
Google releases Gemma 4 QAT. โจ
You can now run Gemma 4 at 3x less memory with near original performance.
QAT makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs: https://huggingface.co/collections/unsloth/gemma-4-qat
QAT Guide: https://unsloth.ai/docs/models/gemma-4/qat
You can now run Gemma 4 at 3x less memory with near original performance.
QAT makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs: https://huggingface.co/collections/unsloth/gemma-4-qat
QAT Guide: https://unsloth.ai/docs/models/gemma-4/qat
posted an update about 2 months ago
Post
9323
Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs.
Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.
You can run and train the model via Unsloth Studio.
GGUF: unsloth/gemma-4-12b-it-GGUF
Guide: https://unsloth.ai/docs/models/gemma-4
Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.
You can run and train the model via Unsloth Studio.
GGUF: unsloth/gemma-4-12b-it-GGUF
Guide: https://unsloth.ai/docs/models/gemma-4
posted an update 2 months ago
Post
2831
Qwen3.6 MTP is here! Run locally on 20GB RAM. โก๏ธ
MTP enables Qwen3.6 to generate ~1.4โ2.2ร faster with no accuracy change.
Qwen3.6-27B: unsloth/Qwen3.6-27B-MTP-GGUF
Qwen3.6-35B-A3B: unsloth/Qwen3.6-35B-A3B-MTP-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.6#mtp-guide
MTP enables Qwen3.6 to generate ~1.4โ2.2ร faster with no accuracy change.
Qwen3.6-27B: unsloth/Qwen3.6-27B-MTP-GGUF
Qwen3.6-35B-A3B: unsloth/Qwen3.6-35B-A3B-MTP-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.6#mtp-guide
posted an update 2 months ago
Post
5991
Weโre excited to announce that Unsloth has joined the PyTorch Ecosystem! ๐ฅ๐ฆฅ
Unsloth is an open-source project that makes training & running models more accurate and faster with less compute. Our mission is to make local AI accessible to everyone. Thanks to all of you for making this possible! ๐
Blog: https://unsloth.ai/blog/pytorch
GitHub: https://github.com/unslothai/unsloth
Unsloth is an open-source project that makes training & running models more accurate and faster with less compute. Our mission is to make local AI accessible to everyone. Thanks to all of you for making this possible! ๐
Blog: https://unsloth.ai/blog/pytorch
GitHub: https://github.com/unslothai/unsloth
posted an update 2 months ago
Post
7788
We collaborated with NVIDIA to teach you how we made LLM training ~25% faster! ๐
Learn how 3 optimizations help your home GPU train models faster:
1. Packed-sequence metadata caching
2. Double-buffered checkpoint reloads
3. Faster MoE routing
Guide: https://unsloth.ai/blog/nvidia-collab
GitHub: https://github.com/unslothai/unsloth
Learn how 3 optimizations help your home GPU train models faster:
1. Packed-sequence metadata caching
2. Double-buffered checkpoint reloads
3. Faster MoE routing
Guide: https://unsloth.ai/blog/nvidia-collab
GitHub: https://github.com/unslothai/unsloth
posted an update 3 months ago
Post
8935
We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw.
Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM
Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp
Guide: https://unsloth.ai/docs/basics/api
Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM
Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp
Guide: https://unsloth.ai/docs/basics/api
posted an update 3 months ago
posted an update 3 months ago
Post
5404
Qwen3.6-27B is out now! Run it locally on 18GB RAM. ๐
Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks.
GGUFs to run: unsloth/Qwen3.6-27B-GGUF
Guide + MLX: https://unsloth.ai/docs/models/qwen3.6
Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks.
GGUFs to run: unsloth/Qwen3.6-27B-GGUF
Guide + MLX: https://unsloth.ai/docs/models/qwen3.6
posted an update 3 months ago
Post
2876
Qwen3.6-35B-A3B can now be run locally! ๐
The model is the strongest mid-sized LLM on nearly all benchmarks.
Run on 23GB RAM via Unsloth Dynamic GGUFs.
GGUFs to run: unsloth/Qwen3.6-35B-A3B-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.6
The model is the strongest mid-sized LLM on nearly all benchmarks.
Run on 23GB RAM via Unsloth Dynamic GGUFs.
GGUFs to run: unsloth/Qwen3.6-35B-A3B-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.6
posted an update 3 months ago
Post
5581
You can now fine-tune Gemma 4 for free with our notebooks! ๐ฅ
You just need 8GB VRAM to train Gemma 4 locally!
Unsloth trains Gemma4 1.5x faster with 50% less VRAM.
GitHub: https://github.com/unslothai/unsloth
Guide + Notebooks: https://unsloth.ai/docs/models/gemma-4/train
You just need 8GB VRAM to train Gemma 4 locally!
Unsloth trains Gemma4 1.5x faster with 50% less VRAM.
GitHub: https://github.com/unslothai/unsloth
Guide + Notebooks: https://unsloth.ai/docs/models/gemma-4/train