How to Build Stateful AI Apps with Gemini Interactions API
Learn how to build stateful AI applications with Google's Gemini Interactions API, including conversation state, follow-up requests, structured output, tool calls, background execution, and migration from GenerateContent.
On this page
Google now recommends the Gemini Interactions API as the default interface for new Gemini applications. Unlike a basic single-request model call, the API can maintain server-side conversation state, expose execution steps, run long-running work in the background, and combine models with tools and agentic workflows. The Interactions API became generally available in June 2026, while the older GenerateContent API remains supported but is now considered a legacy interface for new development.Ā
This tutorial builds a small stateful Gemini application in Python. You will create an interaction, continue the conversation with its identifier, inspect the returned steps, and then extend the same pattern toward structured output and background execution. The goal is to show why conversation state changes the way an AI application should be designed rather than simply replacing one API endpoint with another.
Why Use the Gemini Interactions API
A traditional model request usually requires the application to send the relevant conversation history every time it asks the model for another response. The Interactions API can instead store the interaction and let a later request refer to the previous interaction with previous_interaction_id. Google says this server-side state can also improve context-cache reuse in multi-turn conversations, potentially reducing token costs.Ā
The API also provides a common interface for text generation, multimodal understanding, structured outputs, tools, and agentic workflows. That means an application can begin with a simple question-and-answer feature and later add tools or longer-running operations without adopting an entirely different API pattern.
Prepare the Gemini SDK
Google's current Python SDK for the Gemini API is the google-genai package. The current getting-started documentation shows installing or upgrading that package before creating a client.Ā
pip install -U google-genai
Create an API key through Google AI Studio and keep it in an environment variable. Google recommends using an API key to authenticate requests and track usage against the associated account.Ā
export GEMINI_API_KEY="YOUR_API_KEY"
On Windows, create the equivalent environment variable through PowerShell or Windows environment settings. Do not put the key directly into JavaScript that runs in a user's browser. A backend should make the authenticated API request instead.
Make the First Interaction
The current SDK creates a Gemini client and then uses the Interactions API to send the initial input. A simple text request is enough to establish the workflow.
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Explain server-side conversation state in three sentences."
)
print(interaction)
Gemini 3.8 Flash became generally available on September 2, 2026, and Google describes it as a model designed for long-horizon software engineering, autonomous agents, and complex enterprise workflows.Ā
Do not build your application around a hard-coded model choice simply because it works in the first test. Google regularly adds new models and capabilities, so keep the model identifier in configuration where practical.
Read the Model Output
The Interactions API uses a structured response rather than treating the entire response as one plain text field. The API exposes execution information through steps, which can include model output and tool-related events. Google's migration documentation specifically notes that applications should read the new steps structure rather than relying on the older outputs schema.Ā
For a simple text application, you can inspect the returned interaction object and extract the model-output content according to the SDK version you have installed. Keeping this response handling isolated in one function makes future schema changes easier to manage.
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Give me three practical database backup tips."
)
for step in interaction.steps:
print(step)
The important design change is that your application can now observe more than just the final answer. When tools or longer workflows are involved, the execution steps can help your interface display progress and help developers diagnose where an interaction failed.
Continue a Conversation Without Resending History
The main advantage of server-side interaction state appears when the user asks a follow-up question. Instead of manually constructing a new request containing every previous message, pass the identifier of the earlier interaction through previous_interaction_id. Google documents this as one of the core state-management features of the Interactions API.Ā
first = client.interactions.create(
model="gemini-3.8-flash",
input="I am building a Python API for a small website."
)
second = client.interactions.create(
model="gemini-3.8-flash",
previous_interaction_id=first.id,
input="What authentication approach should I consider?"
)
print(second)
The second request can refer to the context established by the first interaction. This is especially useful for applications where conversations become long enough that manually managing every previous message becomes cumbersome.
Understand Stored and Stateless Modes
Server-side state is optional. Google documents a store parameter that can be used to opt out of stored interaction behavior. When an application does not need server-side conversation state, a stateless workflow may be simpler.Ā
The choice should be made deliberately. A customer-support assistant may benefit from persistent interaction state, while a one-off text classifier may have no reason to store an interaction after processing it. Data retention, privacy requirements, debugging needs, and application architecture should all be considered before choosing the default.
Build a Simple Chat Function
Once the basic calls work, wrap them inside a small class. The application can then keep the latest interaction identifier and use it when the user sends another message.
from google import genai
class GeminiChat:
def __init__(self, model="gemini-3.8-flash"):
self.client = genai.Client()
self.model = model
self.previous_id = None
def send(self, message):
kwargs = {
"model": self.model,
"input": message,
}
if self.previous_id:
kwargs["previous_interaction_id"] = self.previous_id
interaction = self.client.interactions.create(**kwargs)
self.previous_id = interaction.id
return interaction
chat = GeminiChat()
chat.send("I need to plan a small Python web application.")
chat.send("Now suggest a simple database structure.")
This pattern is intentionally small. A production application should also associate the stored interaction identifier with the authenticated user or session rather than keeping one global conversation for everyone.
Add System Instructions
An application often needs to establish behavior before the user begins asking questions. For example, a support assistant may need to answer concisely, avoid inventing account information, and ask for clarification when a request is ambiguous. The Interactions API supports the same broader application pattern used for model instructions while providing a unified interface for additional capabilities.
interaction = client.interactions.create(
model="gemini-3.8-flash",
system_instruction=(
"You are a concise technical support assistant. "
"If required information is missing, ask one clear question."
),
input="My application returns an authentication error."
)
Keep instructions focused on behavior that actually matters to the application. Long instruction blocks are harder to maintain and can make debugging difficult when the model behaves differently from what developers expect.
Use Structured Output for Application Data
Free-form text is convenient for humans but often inconvenient for software. If your application needs predictable fields such as a category, priority, and explanation, use structured output instead of trying to extract those values from prose. Google's current Interactions API documentation lists structured output as a supported feature.Ā
A structured response lets the backend validate the model's output before passing it to another component. This is useful for workflows such as ticket classification, product extraction, content metadata generation, or routing an incoming request to the appropriate tool.
response_format = {
"type": "text",
"schema": {
"type": "object",
"properties": {
"category": {"type": "string"},
"priority": {"type": "string"},
"reason": {"type": "string"}
},
"required": ["category", "priority", "reason"]
}
}
The exact schema configuration should follow the current SDK version's response-format interface. Google's May 2026 schema migration changed how structured-output configuration is represented, so copying an older GenerateContent example into a current Interactions API application can produce incompatible requests.
Inspect Tool Calls
The Interactions API is also designed for applications that let a model use tools. Instead of treating a tool call as ordinary text, the application receives it as a distinct execution step. Google documents function calling and server-side tools as part of the Interactions API's supported capabilities.Ā
A basic architecture looks like this:
- User request: The application sends the user's goal to Gemini.
- Tool decision: Gemini determines that a registered function is needed.
- Function execution: Your backend validates and executes the requested operation.
- Tool result: The result is returned to the interaction.
- Final response: Gemini uses the result to produce the next response.
The backend should never blindly execute arbitrary arguments produced by a model. Validate the function name, argument types, permissions, and target resources before allowing a tool call to affect an external system.
Run Long Tasks in the Background
Some AI operations take longer than a normal interactive request. The Interactions API includes background execution for long-running tasks, allowing the application to submit work without keeping the user interface blocked until the operation finishes. Google lists background execution as one of the API's built-in capabilities.
This changes the user-interface design. Instead of showing a spinning request until completion, your application can create the interaction, record its identifier, show a progress state, and retrieve or monitor the resulting interaction later. That pattern is more appropriate for research, large document processing, and other tasks where completion is not instantaneous.
Migrate from GenerateContent Carefully
If you already have an application using generateContent, do not assume that changing one method name is enough. Google's migration documentation describes changes to the response schema, function-call handling, streaming events, stateless history, and structured-output configuration. The older interface remains supported, but new capabilities are being introduced through the Interactions API.
A sensible migration starts with one endpoint rather than rewriting an entire application. First update the SDK, create one equivalent interaction, inspect the new steps response, and then migrate tool calls and structured output. Keep the old implementation available until the new path passes your application's actual tests.
Common Implementation Mistakes
- Sending the entire conversation manually: This defeats much of the convenience of server-side interaction state.
- Reading the old response schema: Current Interactions API responses use the newer steps-based structure.
- Sharing one interaction ID between users: Store state per authenticated user or application session.
- Executing tool calls without validation: Treat model-generated arguments as untrusted input.
- Using background execution for every request: Keep ordinary conversational requests synchronous when they finish quickly.
- Copying old GenerateContent examples: The current Interactions API has different request and response patterns.
Test the State Before Adding More Features
Before adding tools, background jobs, or complex prompts, test whether the conversation state behaves as expected. Start an interaction with a unique piece of information, send a follow-up that depends on that information, and verify that the model can use the earlier context through previous_interaction_id. Then test what happens when the interaction identifier is missing, invalid, or associated with another user's session.
For a production application, also log interaction identifiers, request duration, model identifiers, error categories, and tool execution outcomes without recording sensitive content unnecessarily. Those records make it much easier to determine whether a failure came from the model, your application, a tool, or the interaction-state layer.
What to Build Next
Once the stateful chat loop works, add one narrowly scoped tool. For example, a support assistant could retrieve a product's documentation, summarize the result, and maintain the conversation state for follow-up questions. The Interactions API is designed to grow from basic model calls into multimodal, structured, tool-using, and agentic applications through the same general interface.Ā
The strongest next step is to keep the first tool simple enough that you can verify every execution step. After that foundation is reliable, background execution and additional tools can be added without rebuilding the conversation architecture. Google's current API direction makes the Interactions API the interface to learn for new Gemini application development, so designing around its state and step-based model now can also reduce future migration work.Ā
Written by


