Thank you I fixed it ♥️
Daniel (Unsloth) PRO
AI & ML interests
None yet
Recent Activity
updated a model about 1 hour ago
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF published a model about 1 hour ago
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF updated a model about 1 hour ago
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3BOrganizations
replied to their post about 11 hours ago
posted an update about 22 hours ago
Post
1245
Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on http://unsloth.ai
and GitHub.
GitHub: https://github.com/unslothai/unsloth
Blog and Guide: https://unsloth.ai/docs/desktop
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on http://unsloth.ai
and GitHub.
GitHub: https://github.com/unslothai/unsloth
Blog and Guide: https://unsloth.ai/docs/desktop
posted an update 6 days ago
Post
910
DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️
DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.
DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.
GGUFs: unsloth/DeepSeek-V4-Flash-0731-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.
DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.
GGUFs: unsloth/DeepSeek-V4-Flash-0731-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
posted an update 12 days ago
Post
1585
We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. 🤯
We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...
1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.
GGUF: unsloth/Kimi-K3-GGUF
GitHub repo: https://github.com/unslothai/unsloth
We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...
1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.
GGUF: unsloth/Kimi-K3-GGUF
GitHub repo: https://github.com/unslothai/unsloth
replied to their post 12 days ago
Which ones? We are working on even smaller ones
posted an update 14 days ago
Post
4160
Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.
GGUF: unsloth/Kimi-K3-GGUF
Guide: https://unsloth.ai/docs/models/kimi-k3
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.
GGUF: unsloth/Kimi-K3-GGUF
Guide: https://unsloth.ai/docs/models/kimi-k3
replied to their post 22 days ago
As long as it's above RDNA2 it should work
posted an update 23 days ago
Post
4996
Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on just 3GB VRAM
GitHub: https://github.com/unslothai/unsloth
Blog + Guide: https://unsloth.ai/docs/basics/amd
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on just 3GB VRAM
GitHub: https://github.com/unslothai/unsloth
Blog + Guide: https://unsloth.ai/docs/basics/amd
replied to their post 23 days ago
If you read our graphic, it says you can update the template as well. Most people don't know how to replace the chat template.
replied to their post 24 days ago
Our MLX quants were update: https://huggingface.co/collections/unsloth/gemma-4
replied to their post 24 days ago
It was posted officially by Google: https://x.com/googlegemma/status/2077449152062247219
posted an update 25 days ago
Post
6035
Gemma 4 is now faster and much more accurate! 🚀
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
posted an update 29 days ago
Post
4841
We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU.
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
posted an update about 1 month ago
Post
4378
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4
posted an update about 1 month ago
Post
6343
DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳
Run lossless DeepSeek-V4-Flash on 168GB RAM or
3-bit works on 110GB Mac, RAM, VRAM setups.
Run via Unsloth Studio or llama.cpp.
GGUF: unsloth/DeepSeek-V4-Flash-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
Run lossless DeepSeek-V4-Flash on 168GB RAM or
3-bit works on 110GB Mac, RAM, VRAM setups.
Run via Unsloth Studio or llama.cpp.
GGUF: unsloth/DeepSeek-V4-Flash-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
posted an update about 2 months ago
Post
3393
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5
We gave 3 models the same prompt and compared one-shot outputs.
The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.
Which output do you like best?
GGUF: unsloth/GLM-5.2-GGUF
We gave 3 models the same prompt and compared one-shot outputs.
The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.
Which output do you like best?
GGUF: unsloth/GLM-5.2-GGUF
posted an update about 2 months ago
Post
4627
Google's new DiffusionGemma can now run at 2000+ tokens/sec! ⚡
We made local DiffusionGemma inference 1.8× faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/models/diffusiongemma
We made local DiffusionGemma inference 1.8× faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/models/diffusiongemma
posted an update 2 months ago
Post
1207
Google releases DiffusionGemma.✨
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
Run with 4x faster text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio.
GGUF: unsloth/diffusiongemma-26B-A4B-it-GGUF
Guide: https://unsloth.ai/docs/models/diffusiongemma
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
Run with 4x faster text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio.
GGUF: unsloth/diffusiongemma-26B-A4B-it-GGUF
Guide: https://unsloth.ai/docs/models/diffusiongemma
posted an update 2 months ago
Post
4313
Google releases Gemma 4 QAT. ✨
You can now run Gemma 4 at 3x less memory with near original performance.
QAT makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs: https://huggingface.co/collections/unsloth/gemma-4-qat
QAT Guide: https://unsloth.ai/docs/models/gemma-4/qat
You can now run Gemma 4 at 3x less memory with near original performance.
QAT makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs: https://huggingface.co/collections/unsloth/gemma-4-qat
QAT Guide: https://unsloth.ai/docs/models/gemma-4/qat
posted an update 2 months ago
Post
9375
Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs.
Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.
You can run and train the model via Unsloth Studio.
GGUF: unsloth/gemma-4-12b-it-GGUF
Guide: https://unsloth.ai/docs/models/gemma-4
Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.
You can run and train the model via Unsloth Studio.
GGUF: unsloth/gemma-4-12b-it-GGUF
Guide: https://unsloth.ai/docs/models/gemma-4