Files up to ~1.2 GB
Snappy chat and summarization on most modern phones.
Dhruva loads any quantized GGUF you download from Hugging Face — there's no curated-only walled garden. Below is the starter catalog: known-good starting points with real, verified repo IDs and honest RAM guidance for your phone.
How much phone you need
Files up to ~1.2 GB
Snappy chat and summarization on most modern phones.
Files up to ~3 GB
Stronger reasoning; wants a little headroom.
Files larger than ~3 GB
Flagship devices; the most capable local models.
Starter catalog
| Model | Repo | Role | Quant | Size | RAM floor |
|---|---|---|---|---|---|
| Llama-3.2-1B-InstructChat, summarization — best instruction-following at this size | bartowski/Llama-3.2-1B-Instruct-GGUF | Chat | Q4_K_M | ~770 MB | 4 GB+ |
| Qwen2.5-1.5B-InstructConversational, multilingual, balanced | bartowski/Qwen2.5-1.5B-Instruct-GGUF | Chat | Q4_K_M | ~986 MB | 4 GB+ |
| SmolLM2-1.7B-InstructEfficient, maths-capable — coding and reasoning | bartowski/SmolLM2-1.7B-Instruct-GGUF | Chat | Q4_K_M | ~1.0 GB | 4 GB+ |
| Llama-3.2-3B-InstructBetter reasoning for complex tasks and analysis | bartowski/Llama-3.2-3B-Instruct-GGUF | Chat | Q4_K_M | ~1.9 GB | 6 GB+ |
| Phi-4-mini-instructReliable, well-tuned — research-backed | unsloth/Phi-4-mini-instruct-GGUF | Chat | Q4_K_M | ~2.4 GB | 6 GB+ |
| SmolVLM2-2.2B-InstructPhoto and screenshot understanding | ggml-org/SmolVLM2-2.2B-Instruct-GGUF | Vision | Q4_K_M + mmproj-Q8_0 | ~1.6 GB total | 6 GB+ |
| All-MiniLM-L6-v2Document search and RAG — always paired with a chat model | second-state/All-MiniLM-L6-v2-Embedding-GGUF | GGUF | ~20 MB | negligible |
Sizes are approximate Q4_K_M file sizes at time of writing — always check the exact size on the model's Hugging Face page before downloading, especially on a metered connection.
Before you download
Open a model in the hub and Dhruva shows a verdict badge instead of making you guess. The logic is deliberately simple and file-size-based, not a black box.
Dhruva reads the quantized file size of the GGUF and buckets it: up to 1.2 GB is the "1B class", up to 3 GB the "3B class", larger is "4B+".
Each bucket has a RAM floor: 4 GB for 1B, 6 GB for 3B, 8 GB for 4B+.
Your device's RAM is compared to that floor. Below it → Not recommended. At the floor but under 1.5× → Possible (little headroom, expect it slower). At 1.5× or above → Comfortable.
File size, not parameter count, drives the bucket — a Q4_K_M quant tracks real memory pressure far more closely than an advertised parameter count. Vision models include their mmproj file in the bucketed size, since it loads alongside the base model.
What the quant name means
Quantization shrinks a model's weights from 16-bit floats to a few bits each — a fraction of the size and memory, for a small, usually unnoticeable quality cost. It's what makes running a capable model on a phone possible at all.
The sweet spot Dhruva recommends: ~4-bit, K-quant medium. Best size-to-quality trade for phones.
A step up in fidelity for a bigger file. Worth it when you have the RAM to spare.
Near-lossless. Used for the small vision projector that rides alongside a base model.