Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally
Today, we’re excited to launch Unsloth Desktop, our new open-source desktop app for running and training models locally on Windows, macOS and Linux.
Unsloth Desktop is easy to download and requires no setup. It’s available for nearly all modern CPU and GPU setups and most operating systems.
Here is an overview of everything Unsloth supports:
Supports text, MLX, diffusion image/video, audio, GGUF and more models
Self-healing tool calls are 50% more accurate, with sandboxed Python and Bash execution
Image and video generation + training, with faster diffusion inference
Day-zero support for Qwen3.8, Muse Glimmer 30B, MiniMax H3, Kimi K3, DeepSeek-V4 Flash and Gemma 4
Unlimited private web search, Deep Research, RAG and MCP
Parallel chats, each with its own tools, sandbox and prompt queue
An OpenAI-compatible API, plus connections to cloud OpenAI and Anthropic models
Secure remote access from any device via Cloudflare HTTPS
Runs on NVIDIA, AMD, Intel, Apple Silicon, CPU and multi-GPU systems
Image + video generation 🎬
You can now generate and train images and videos locally, right inside Desktop. Diffusion is inference faster, LoRA workflows are supported, and video generation works on Apple Silicon through Metal and via CPU setups.
MiniMax H3 is supported and it generates video with synchronized audio, and takes images, videos or audio as references. On an NVIDIA B200, a 960×544, 124-frame, 8-step generation dropped from 70+ seconds to around 13 seconds.
Better tool calling 🔧
Tool calling is now 50% more accurate with self-healing tool calls. Your local models can run Python and Bash in a sandbox to calculate, analyze data, test code and generate files, and malformed tool calls get repaired and retried automatically. You can download any files or artifacts your chats produce.
You can also connect your local model to most agents. Run:
unsloth start claudeto connect your local model with Claude Code. You can switch between Claude Code, Codex, and other agents interchangeably.
Deep Research + Local RAG 🔎
Desktop includes unlimited private web search, deep research, RAG and MCP. Deep Research builds a plan, reads sources, tracks citations and produces a full report, and you can combine it with your own knowledge bases.
You can also link a local folder directly to RAG and it stays synced. Add, rename or update files and Unsloth handles it.
You can also start a new conversation while another model is still generating and do parallel chats.
Meta Muse Glimmer 30B
Today’s also comes with support for Meta’s new dense 30B coding model ‘Muse Glimmer’ which runs locally on around 18GB RAM/VRAM setups, and you can fine-tune the model with Unsloth.
Run Local Models Anywhere 💻
Unsloth Desktop works on NVIDIA, AMD, Intel/Vulkan, Apple Silicon, CPU and multi-GPU systems, with GPU layer controls, memory offloading and tensor parallelism for GGUF models. You can serve anything through the OpenAI-compatible API, expose it securely over Cloudflare HTTPS, and essentially access the model remotely via a phone or another device.
Unsloth Desktop GitHub
Unsloth Desktop is available now for Windows, macOS and Linux.
You can also download Unsloth Desktop from our GitHub releases page. We’d love for you to try it out and are more than happy for any feedback!
A huge thanks to NVIDIA and Hugging Face for collaborating with us on this release! And most importantly, we’re grateful for our early alpha testers and couldn’t have done it without them.
We hope to share more exciting news soon, especially around Qwen3.8.




Nice tool