Skip to content

How to Run Gemma 4 Locally in Android Studio Quail 4

Android Studio Quail 4 can run Gemma 4 locally for private, offline AI coding. Learn how to install the IDE, download the right model, use Agent Mode, and review local AI changes safely.

How to Run Gemma 4 Locally in Android Studio Quail 4

On this page

Android Studio Quail 4 can now run Gemma 4 directly on your computer, giving developers a way to use an AI coding agent without sending source code to a cloud model. The stable Quail 4 release includes a built-in runtime for Gemma 4, one-click model management, and support for agentic coding tasks such as multi-file refactoring. This guide shows how to run Gemma 4 locally, choose the right model for your hardware, and verify that your coding requests are actually being processed locally.

What changed in Android Studio Quail 4

Android Studio Quail 4, version 2026.1.4, is the final stable release in the Quail series. Its Gemma 4 integration is different from simply connecting Android Studio to a cloud provider because the IDE includes the runtime needed to execute the local model. You can download the model from Android Studio itself and select it for Agent Mode without installing a separate model server. This makes local AI coding practical for developers who work with sensitive source code, restricted networks, or projects where cloud token limits are inconvenient.

Google also bundled 23 Android-specific skills into Quail 4. These are curated instructions that help the coding agent handle Android APIs, migrations, profiling, navigation, adaptive layouts, and other platform-specific tasks. The important distinction is that Gemma 4 provides the model while Android skills provide additional Android-focused guidance. Using both can make a local coding workflow more useful than asking a general-purpose local model to infer every Android convention from scratch.

Check your hardware before downloading Gemma 4

Local inference moves the computing cost from a remote server to your own machine, so hardware matters. Google documents two practical Gemma 4 choices for Android Studio: Gemma E4B requires 12 GB of total system RAM and about 4 GB of storage, while Gemma 26B MoE requires 24 GB of total system RAM and about 17 GB of storage. Total RAM includes Android Studio and the rest of the development environment, so the minimum figure should not be treated as a comfortable recommendation for large projects.

If your computer has 12 GB of RAM, start with the smaller model rather than trying to force the larger model into memory. A machine with 24 GB can meet the documented requirement for the 26B model, but Android Studio, Gradle, an emulator, a browser, and other development tools can consume substantial memory at the same time. Google says machines with 32 GB or more will run the broader local-model experience more comfortably. Storage is also worth checking before installation because model files are much larger than a typical Android Studio plugin.

Gemma 4 modelTotal RAMStoragePractical choice
Gemma E4B12 GB4 GBLower-spec development machines
Gemma 26B MoE24 GB17 GBHigher-performance development machines

Install Android Studio Quail 4 first

Before configuring Gemma 4, make sure Android Studio is running Quail 4 rather than an older Quail build. Open Android Studio and check for updates if you already have the IDE installed. The stable Quail 4 release is identified as version 2026.1.4, and Google lists it as the current stable release in the Quail series.

  1. Open Android Studio.
  2. Open the Help menu on Windows or Linux, or the Android Studio menu on macOS.
  3. Choose the option for checking updates.
  4. Install Android Studio Quail 4 if an update is available.
  5. Restart Android Studio after the installation finishes.

The result you want is Android Studio Quail 4 running normally with your existing project available. If your installation still reports an older version, finish the update before trying to configure Gemma 4. This avoids looking for model settings that are unavailable in older builds.

Open the Gemma model provider settings

Once Quail 4 is installed, the local model configuration is handled inside Android Studio. You do not need to install Ollama or another third-party inference server for the native Gemma 4 integration. Android Studio manages the model download and uses its integrated inference runtime.

  1. Open Android Studio Settings.
  2. Open Tools.
  3. Select AI.
  4. Open Model Providers.
  5. Select Gemma.

You should now see the Gemma provider and the available model choices. This is the key difference from older local-model workflows, where developers often had to install a separate runtime and then connect Android Studio to it. Quail 4 handles the native Gemma path from the IDE, reducing the number of components you have to maintain.

Download and select the Gemma 4 model

Choose the Gemma model that fits your available memory and storage. Android Studio handles downloading, verifying, and updating the model weights. If you are working on a machine close to the minimum requirements, the smaller E4B option is the safer starting point because leaving enough memory for Android Studio and Gradle is more useful than selecting a larger model that causes constant memory pressure.

  1. Select the Gemma model provider.
  2. Choose the Gemma 4 model you want to use.
  3. Start the model download.
  4. Wait for Android Studio to finish verifying the model.
  5. Open the Agent model selector after the download completes.
  6. Select Gemma 4 as the active model.

The model is ready when Android Studio allows you to use it from its AI coding workflow. The first request can take longer than later requests because your machine has to load the local model and allocate the required memory. That delay is expected and is different from a cloud API request, where the model is already running on a remote server.

Run your first local coding task

Do not begin by asking Gemma 4 to redesign your entire application. Start with a small task that gives you an easy way to verify both the model and the agent workflow. For example, open a small Android project and ask the agent to identify a duplicated utility function and suggest a safe refactoring. A local agent can inspect the project and, where supported, use its tool-calling capabilities to work through multiple files.

  1. Open an Android project in Android Studio.
  2. Open Agent Mode.
  3. Confirm that Gemma 4 is selected as the active model.
  4. Ask the agent to inspect one small, clearly defined coding problem.
  5. Review the proposed changes before applying them.
  6. Build the project after accepting the changes.

A successful test should give you a response from the selected local model and leave the project in a state that can be checked with the normal Android build process. Treat the generated changes like code from any other AI assistant: inspect the diff, check API usage, run tests, and verify the build. Local execution improves where the data is processed, but it does not automatically make generated code correct.

Use Gemma 4 for multi-file work carefully

Gemma 4 supports agentic tool calling, which means the model can work through a sequence of tool operations instead of producing only a single block of suggested code. In Android Studio, this can be useful for tasks such as coordinated refactoring across several files. For example, a request to rename a shared class can require changes to its declaration, imports, references, and related tests.

The useful part is that the agent can reason about the task as a workflow rather than treating every file independently. The risk is also clear: a multi-file operation has a larger blast radius than changing one function. Review the generated diff before accepting it, especially when the project contains build configuration, database migrations, authentication logic, or release-critical code.

Understand what stays local and what does not

When you select the native Gemma 4 local model, inference runs on your development computer rather than through a third-party cloud model provider. That means source code used by this local inference path can remain on the machine, which is useful for organizations with strict source-code handling policies. It also means requests processed locally do not depend on a cloud model's token quota.

That privacy advantage should not be generalized to every AI feature inside Android Studio. Quail 4 also supports cloud-based Gemini and other model providers, and those workflows have different data paths. Before working with proprietary code, check which model is selected and which provider is handling the request. The model name alone is not enough to determine where every AI-assisted operation is processed.

Know the trade-off before replacing cloud AI

Local Gemma 4 is not simply a free replacement for the strongest cloud model. Google's own documentation warns that local models can provide lower performance, higher latency, and more limited feature support than cloud-based Gemini models. Your actual experience also depends heavily on CPU, memory, storage speed, and the size of the project being analyzed.

The best workflow may therefore be a hybrid one. Use Gemma 4 when source-code privacy, offline work, or predictable local execution matters, and use a cloud model when a task benefits from higher capability or larger remote resources. Android Studio's model selection makes that choice part of the development workflow instead of forcing every project into one provider.

Check the result before trusting the agent

After Gemma 4 completes a task, review the changes as normal production code. Look for incorrect Android APIs, unnecessary dependencies, lifecycle mistakes, broken imports, and changes that solve the immediate request while creating a different problem elsewhere. Run the project's build and tests rather than treating a successful AI response as proof that the implementation works.

For larger refactoring tasks, make smaller commits so you can identify which AI-generated change introduced a regression. This is particularly useful with local agent workflows because the absence of cloud limits can make it tempting to give the model increasingly broad instructions. Keeping the task boundary clear makes both the generated changes and the review process easier to control.

When Gemma 4 local coding makes the most sense

Gemma 4 is most compelling when keeping code on the development machine is a real requirement rather than simply a preference. It is also useful for developers who regularly work without reliable internet access or who want to avoid depending on cloud usage quotas for routine coding assistance. Quail 4 removes much of the setup friction that previously came with local AI coding by putting model management and inference support inside the IDE.

The practical starting point is simple: install Quail 4, check your available RAM and storage, download the smallest model that fits comfortably, select Gemma 4 in Agent Mode, and test it on a small project task. Once you know how your hardware performs, you can decide whether the larger model is worth the additional memory and storage. That measured approach is more useful than choosing a model by size alone, because local AI performance is ultimately constrained by the machine sitting beside your keyboard.

S

Written by

Sarah Khan

I’m fascinated by artificial intelligence and the rapid changes happening around AI tools, models, and agents. I enjoy testing new AI technologies, following important developments, and understanding how they can be useful in real life. I like explaining complex AI topics in a simple and practical way.

51 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.