MOONLIGHT AI

Private AI.
Running on your device.

Run capable AI models locally, keep conversations on-device, and continue working without an internet connection after your models are installed.

Coming soon to Google Play
See how it works
9:41
M
Phi-3 Mini
Running locally ? 19.4 t/s
Can you explain how local LLM inference works without the internet?
M

Local LLM inference runs on your device's hardware (CPU/GPU) rather than sending data to a cloud server.

Moonlight uses llama.cpp which is highly optimized for mobile devices. The model weights are stored locally on your phone's storage in a quantized format (GGUF).

Because the computation happens on-device, it works offline and your chat history is stored locally.

Message Moonlight...
Step 01

Model Downloaded

GGUF weights retrieved from Hugging Face.

Step 02

Integrity Verified

SHA-256 checksums confirmed on local storage.

Step 03

Model Installed

Ready for execution on the neural engine.

Step 04

Loaded into Memory

llama.cpp maps weights into device RAM.

Step 05

Prompt Enters

Your text is tokenized locally.

Step 06

Local Inference Begins

CPU/GPU calculates probabilities.

Step 07

Tokens Stream

De-tokenization generates readable text.

Step 08

Watch intelligence stay local.

Nothing about this required a cloud server.

Phi-3 Mini
9:41
M
Phi-3 Mini
What happens to my data?
M
Privacy by Architecture: Chats stay on your device. It never touches a network.
PRIVACY BY ARCHITECTURE

Your AI stays close.

Moonlight is designed around a fundamental boundary. What happens on your device, stays on your device.

Internet
Network Required
Model Download
Play Store Access
Your Device
Strictly Local Boundary
GGUF Model
Stored in local storage
Inference Engine
llama.cpp • CPU/GPU
Private Conversations
Stored in local database
Available Now

Meet Moonlight AI.

A different approach to personal AI — on your device, under your control.