Hermes 2 Pro Llama 3 8B
General-purpose Llama 3 8B finetune for function calling and JSON
Every open-source model that fits an iPhone's memory budget inside Private LLM — running fully offline, with nothing sent to any server.
88 models
General-purpose Llama 3 8B finetune for function calling and JSON
Uncensored Llama 3 8B fine-tune lacking built-in refusals
Uncensored instruction-tuned Llama 3.1 8B with user-defined boundaries via system prompt
Reasoning with chain-of-thought, distilled from DeepSeek R1 into Llama 3.1 8B
Uncensored DeepSeek R1 Distill 8B via refusal abliteration
General-purpose assistant boosting instruction following over Llama 3.1 8B
Roleplay specialist fine-tuned from Llama 3 8B Instruct
General assistant with function calling, structured JSON, and ChatML (Llama 3 8B)
General-purpose Llama 3.1 8B model handling 131k-token contexts and function calling
Roleplay model for trauma and self-harm narratives on Llama 3 8B
Uncensored behavioral reversal of Llama 3 8B Instruct using weight orthogonalization
General instruction following aligned via self-play preference optimization on Llama 3 8B
General finetune of Llama 3 8B with improved first-turn instruction following
Vulnerability explanation and security code generation specialist on Llama 3 8B
Uncensored Llama 3.1 8B finetune with safety refusals removed
Medical exam and clinical knowledge specialist built on Llama 3.1 8B
General conversational assistant surpassing GPT-3.5 turbo in standard benchmarks
General-purpose model in the Llama 3 8B family
Uncensored Llama 3 8B variant with refusal direction removed
General-purpose model in the Llama 3.1 8B family
Uncensored Llama 3.1 8B Instruct with refusal disabled by abliteration
Survival advisor finetuned from Llama 3.1 for step-by-step shelter building
Uncensored safety-ablated Llama 3 8B fine-tuned with DPO
General chat and coding expert finetuned from Llama 3 8B
Survival biomedical finetuned from Llama 3 8B with 8192 context
Step-by-step reasoning distilled from DeepSeek R1 into Qwen 7B
Roleplay finetune of Qwen 2.5 7B for adaptive persona and storytelling
General assistant merging knowledge from larger models into Qwen 2.5
Solves GitHub issues and refactors code, built on Qwen 2.5
General assistant skilled in coding and math reasoning
Code generation and refactoring instruction-tuned on Qwen 2.5 with 32K context
Multilingual assistant with top Japanese scores, Mistral 7B based
Hebrew-first multilingual chat and translation model based on Mistral 7B
General instruction-tuned assistant skilled in function calling
General conversation model trained for accurate function calling and JSON generation
Code completion and refactoring specialist built on OpenChat 3.5
Uncensored Mistral 7B model requiring user-applied safety filters
Generalist chat, reasoning, math, and coding model built on Mistral 7B
Generalist trained on one million GPT-4 generated instructions built on Mistral 7B
General conversational and coding model trained with RLHF on Mistral 7B
General instruction-following model for long documents and conversations, Mistral 7B
Uncensored Llama 2 finetune de-aligned for NSFW and unrestricted generation
Instruction following specialist built on Llama 2 7B
Multilingual dialogue and reasoning model fine-tuned from Yi 6B
Uncensored abliteration of Qwen3 4B for unrestricted research
General assistant with 262K context, improved reasoning, and creative writing
Uncensored Qwen3 4B with 262K context and sharply reduced refusals
Uncensored model in the Qwen3 4B family
Uncensored Qwen3 4B abliteration that reduces refusals across 262K tokens
Uncensored Phi 3 Mini 3.8B with refusal safeguards largely stripped
General instruction-following Phi 3 Mini 3.8B with math and logic proficiency
Uncensored via refusal ablation on Llama 3.2 3B Instruct
General assistant fusing source models for improved instruction following
General-purpose assistant with reliable function calling and JSON outputs
General-purpose model in the Llama 3.2 3B family
Uncensored Qwen 2.5 model with full system prompt steerability
General knowledge model with strong structured output and JSON capabilities
Coding assistant handling generation, reasoning, and fixes across long files
Uncensored Llama 3.2 3B instruct with user-defined alignment and 131K context
Uncensored Llama 3.2 3B modification that bypasses refusal mechanisms
Complex summaries of advanced topics from an instruction-tuned StableLM 3B assistant
Detailed general assistant finetuned from StableLM 3B with DPO
Uncensored conversation, roleplay, and agent tasks from a Phi 2 base
General instruction following model from Phi 2 with reliable commonsense reasoning
Instruction-following conversational finetune of Phi 2 3B with simple chat template
General-purpose fine-tune of Gemma 2 2B with multilingual strengths
General chat assistant fine-tuned for instruction following and factual accuracy
Roleplay-focused distillation of Qwen 2.5 with 131K context
General instruction assistant from StableLM 3B outperforming 10x larger models
Uncensored instruction-tuned chat with user-defined alignment on Qwen 2.5 1.5B
Instruction-tuned general assistant reliably producing clean JSON
Code completion, refactoring, and explanation expert from Qwen 2.5
Uncensored model derived by abliterating Llama 3.2 1B Instruct
General compact assistant with sharp instruction following from multi-model distillation
General-purpose model in the Llama 3.2 1B family
Uncensored conversation specialist built on TinyLlama 1.1B
Conversational assistant fine-tuned from TinyLlama 1B for on-device use
Uncensored instruct finetune of Llama 3.2 1B for coding and agentic tasks
Uncensored Gemma 3 1B finetune for neutral factual responses
Uncensored Gemma 3 1B via abliteration
Uncensored Qwen 2.5 0.5B instruct with full system prompt control
General-purpose 0.5B model with strong instruction following and JSON output
Code generation, reasoning, and repair via Qwen 2.5 with 32K context
On iPhone, Private LLM runs models up to Hermes 2 Pro Llama 3 8B, depending on how much memory your specific iPhone has.
Private LLM runs 88 models on iPhone, all fully on-device.
Yes. Once downloaded, every model runs 100% on-device on your iPhone — no internet connection and no data sharing.