on-device AI, working together
Several small AI models team up right on your iPhone - a router picks the specialist, replies stream instantly, and nothing you type ever leaves the device.
noni is a team of small AI models that live on your iPhone.
Instead of one giant model in the cloud, noni orchestrates several small open-weight language models locally: a tiny router reads each request and hands it to the right specialist - general chat, coding, or writing. Nothing you type ever leaves your phone.
PRIVATE BY ARCHITECTURE
• zero data collection - no account, no analytics, no tracking
• conversations are stored only on your device
• the network is used for one thing: downloading model weights from Hugging Face
• after a model is downloaded, it works in airplane mode
A TEAM, NOT A SINGLE MODEL
• auto - a 0.6B router classifies each request and picks the specialist
• solo - pin any downloaded model and talk to it directly
• relay - pipelines like draft → polish, one model refining another's work
• council - two models answer, a third merges the best of both
BUILT FOR IPHONE
• Apple's MLX framework with 4-bit quantized models and streaming replies
• markdown rendering with copyable code cards
• background, resumable model downloads
• automatic conversation titles, edit & resend, regenerate
• a strict black-and-white design: no clutter, no color, no distractions
MODELS & SIZES (download only what you need)
• qwen3 0.6b - router and quick chat, ~350 MB
• llama 3.2 1b - everyday general model, ~710 MB
• gemma 3 1b - rewriting and summarizing, ~770 MB
• qwen2.5 coder 1.5b - writes and explains code, ~880 MB
• qwen3 1.7b - stronger generalist for recent iPhones, ~980 MB
• llama 3.2 3b - the most capable, for high-memory iPhones, ~1.8 GB
The starter pack (router + general model) is about 1.1 GB. Models are free, open-weight, and hosted on Hugging Face.
HONEST EXPECTATIONS
Small local models are fast, private, and surprisingly capable for everyday questions, quick code snippets, and rewrites - but they are not cloud giants. Expect roughly 10-30 tokens per second on recent iPhones, and occasional mistakes. That is the trade for total privacy, zero cost, and offline use.
Requires an iPhone with iOS 17 or later. On older or lower-memory devices, stick to the smaller models; the library labels which models want a high-memory iPhone.
Built with Llama. Qwen models are Apache-2.0 by Alibaba Cloud; Gemma is provided under Google's Gemma Terms of Use; inference runs on Apple's open-source MLX.
Chrome-Stats does not own this Apple app. Please use these information below to contact the Apple app developer.