Skip to main content
· 4 min readAIRN
Malik Chohra

By Malik Chohra

Running Llama on iPhone and Android with React Native

How to run a local LLM directly on-device using React Native and llama.rn. A practical guide to memory constraints, model quantization, and offline AI architecture.

Running a quantized Llama-family model directly on iOS and Android is practical for privacy-first apps. In this tutorial, we show how to run a local GGUF model through llama.rn, completely offline, while navigating the strict memory limits of mobile operating systems.

💡 Want to skip the native configuration?

AI Mobile Launcher AI Pro ships on-device inference through llama.rn. Start from the pre-integrated feature instead of wiring the native bridge from scratch; you still need a development build for native modules.

Why Run Llama 3 On-Device? (The Privacy Mandate)

Cloud APIs like GPT-4o are incredibly powerful, but they are fundamentally disqualified for certain use cases. If you're building an app that processes user medical records, local encryption layers aren't enough, the data simply cannot leave the device.

On-device AI guarantees zero latency variations, complete offline availability (perfect for travel or field-work apps), and structural compliance with HIPAA and GDPR data residency strictures. Running Meta's Llama 3 parameters entirely inside the mobile CPU/NPU bridges the gap between current AI and absolute data sovereignty.

The Mobile Memory Constraint: Why Quantization is Mandatory

The raw Llama 3 8B model requires roughly 16GB of VRAM to run effectively. A standard iPhone 13 has 4GB of total RAM, and iOS aggressively terminates background apps that consume more than 2GB.

To fit an 8-billion parameter model into an iPhone, we must mathematically compress it through a process called Quantization.

What is 4-bit Quantization?

Quantization reduces the precision of the model's weights from 16-bit floats to 4-bit integers. This shrinks the model file from 16GB down to roughly 4.3GB. While 4.3GB is still heavy for a mobile app bundle, we can load it efficiently into the device's Neural Engine (NPU) using optimized C++ bindings like `llama.cpp`.

Setting Up LLM C++ Bindings in React Native

We cannot run Python local servers inside a mobile app. Instead, we use `llama.cpp` wrapped inside a React Native JSI (JavaScript Interface) boundary. This allows our JavaScript code to synchronously communicate with the highly optimized C++ inference engine.

While you can build the native bridge from scratch, llama.rn handles the native inference surface and thread management for React Native.

// Example of initializing an on-device Llama context
import { initLlama } from 'llama.rn';

// Inside your initialization hook
const loadLocalModel = async () => {
  try {
    const context = await initLlama({
      model: 'models/llama-3-8b-instruct-q4_k_m.gguf',
      use_mlock: true, // Keep model loaded in RAM
      n_ctx: 1024, // Limit context to save memory
      n_threads: 4     // Optimized for mobile CPUs
    });
    
    console.log("On-device model loaded successfully!");
    return context;
  } catch (error) {
    console.error("Device ran out of memory.", error);
  }
};

Building a Privacy-First App?

Configuring C++ bindings, memory locks, and native operators in React Native is complex. Let our engineering team handle the infrastructure while you focus on the product.

Contact us for Custom Development →

Handling iOS and Android Background Termination

The biggest hurdle in on-device AI isn't computational speed, it's the OS memory watchdog. If your App uses 1.5GB of RAM to infer a prompt and the user switches to the Camera app, iOS will kill your app instantly.

To mitigate out-of-memory (OOM) crashes:

  • Only load the `.gguf` model file precisely when the user navigates into the AI chat interface.
  • Use `AppState.addEventListener` to explicitly eject the model from RAM (`context.release()`) the millisecond the app goes into the background.
  • Restrict the context window to exactly what you need. A 4096 context window consumes drastically more working memory than a 1024 context window.

Streaming Responses from the Native Thread

Just like calling cloud providers, on-device inference takes time. Generating a single token on an iPhone 15 Pro takes roughly 25-40ms. To maintain a smooth UI, we must stream the tokens from the C++ layer over the React Native bridge.

// Streaming the on-device inference to the UI
const generateResponse = async (prompt) => {
  setStreamingText("");
  
  await llamaContext.completion(
    { prompt, n_predict: 200 },
    // Callback fires as the C++ engine yields tokens
    (chunk) => {
      setStreamingText((prev) => prev + chunk.token);
    }
  );
};

Using this pattern, the user perceives the app as instantly responsive, even though their phone's processor is working near 100% capacity to generate the text.

The Reality of Battery Drain and Thermal Throttling

We must address the elephant in the room. Running Llama 3 locally will drain a modern smartphone battery by approximately 1% for every 2-3 minutes of active inference.

Additionally, heavily utilizing the Neural Engine generates significant heat. After 10 consecutive minutes of generating text, the operating system will begin CPU thermal throttling, slowing token generation by up to 50%. On-device AI is best used for burst-tasks (e.g., summarizing an offline note, extracting structured JSON from a photo) rather than prolonged, 30-minute chat sessions.

Summary

  • On-device LLMs provide absolute privacy and zero-latency offline availability.
  • Models must be compressed using 4-bit Quantization (GGUF format) to fit inside 2GB of mobile RAM.
  • Use JSI and C++ bindings (`llama.cpp`) to run inference without the overhead of the JavaScript bridge.
  • Aggressively manage memory drops during App backgrounding to avoid OS termination.

Need help building offline AI capabilities?

We specialize in deeply-integrated, private LLM deployments for mobile scale.

Tell Us About Your Project

Related Articles