GitHub Copilot SDK Python Tutorial: Build a Tool-Using AI Agent
Build a Python AI agent with the GitHub Copilot SDK, add a custom tool, stream responses, and learn how to handle permissions and persistent sessions.
On this page
The GitHub Copilot SDK now gives Python developers a direct way to put Copilot's agent runtime inside their own applications. Instead of building the planning loop, tool-calling layer, and session handling yourself, you can create a Copilot client, start a session, give it a Python function as a tool, and let the agent decide when that function is useful. GitHub currently documents Python 3.11+ support and provides the SDK as the github-copilot-sdk package.
This tutorial builds a small command-line AI agent in Python and then gives it a custom calculator tool. The important part is not the calculator itself; it is the pattern. Once you understand how a session and custom tool fit together, you can replace the example function with an internal database lookup, ticket system, file operation, or another controlled application capability.
Why use the Copilot SDK instead of calling a model directly?
A conventional model API usually gives your application a model endpoint and leaves the surrounding agent architecture to you. The GitHub Copilot SDK takes a different approach: it exposes the same underlying agent runtime used by Copilot CLI through SDKs for Python, TypeScript, Go, .NET, Java, and Rust. GitHub says the runtime can handle planning, tool invocation, file edits, and related agent behavior programmatically.
That makes the SDK useful when your application needs more than a single prompt and response. A session gives the agent a place to work, while tools let it call application code when a task requires an operation that the language model cannot perform by itself. The SDK also supports streaming events, session history, asynchronous Python code, and multiple transport options.
This is also where the approach differs from building another standalone Copilot CLI workflow. Here, the agent becomes a component inside your Python application rather than something the user operates only through the terminal.
What you need before writing the agent
You need Python 3.11 or newer and access to GitHub Copilot. GitHub's current quickstart also supports Node.js, Go, Rust, Java, and .NET, but this walkthrough stays with Python. The Python package is installed with pip install github-copilot-sdk. The SDK can manage its runtime automatically, and GitHub documents a separate runtime-download command when you want to pre-provision it.
For a local experiment, create a clean project directory and install the package:
mkdir copilot-python-agent
cd copilot-python-agent
python -m venv .venvActivate the virtual environment using the command appropriate for your operating system, then install the SDK:
pip install github-copilot-sdkThe current Python SDK documentation also supports downloading the managed runtime explicitly with python -m copilot download-runtime. If you skip that step, the SDK can stage the required runtime automatically when it first needs it.
Create the first Copilot session
Start with the smallest useful program. Create a file named main.py and initialize a CopilotClient. A session is the conversation and execution context in which the agent receives instructions and produces work.
import asyncio
from copilot import CopilotClient
from copilot.session import PermissionHandler
async def main():
client = CopilotClient()
await client.start()
```
session = await client.create_session(
on_permission_request=PermissionHandler.approve_all,
model="auto",
)
response = await session.send_and_wait(
"Explain what an AI agent does in two sentences."
)
print(response.data.content)
await client.stop()
```
asyncio.run(main())The important sequence is deliberately short: start the client, create a session, send a prompt, read the response, and stop the client. GitHub's current Python quickstart follows the same basic lifecycle. The auto model setting lets Copilot select from the models available to the runtime rather than hard-coding one model name.
The example uses PermissionHandler.approve_all to keep the tutorial simple. That is a convenience setting for an example, not a permission strategy you should automatically carry into a production application. An application that exposes file changes, shell operations, or other consequential tools should decide explicitly which operations the agent may perform.
Give the agent a tool it can actually call
An AI agent becomes more useful when it can hand part of a task to ordinary application code. The SDK's Python package supports custom tools and can generate their input schema from Pydantic models. In practical terms, the model decides that it needs a tool, the SDK validates the arguments, your Python function runs, and the result goes back into the agent's context.
For a first example, use a calculator rather than an external service. That keeps the tutorial reproducible while demonstrating the same mechanism you would use for a real application integration.
import asyncio
from pydantic import BaseModel, Field
from copilot import CopilotClient, define_tool
from copilot.session import PermissionHandler
class CalculatorParams(BaseModel):
first: float = Field(description="The first number")
second: float = Field(description="The second number")
operation: str = Field(
description="The operation: add, subtract, multiply, or divide"
)
@define_tool(description="Perform a basic arithmetic calculation")
async def calculator(params: CalculatorParams) -> str:
if params.operation == "add":
result = params.first + params.second
elif params.operation == "subtract":
result = params.first - params.second
elif params.operation == "multiply":
result = params.first * params.second
elif params.operation == "divide":
if params.second == 0:
return "Cannot divide by zero."
result = params.first / params.second
else:
return "Unsupported operation."
```
return str(result)
```
async def main():
client = CopilotClient()
await client.start()
```
session = await client.create_session(
on_permission_request=PermissionHandler.approve_all,
model="auto",
tools=[calculator],
)
response = await session.send_and_wait(
"Calculate 48 multiplied by 17, then tell me the result."
)
print(response.data.content)
await client.stop()
```
asyncio.run(main())There are three pieces worth understanding. CalculatorParams describes the arguments the tool accepts, @define_tool exposes the function to Copilot, and tools=[calculator] makes that function available to the session. The model is not being given permission to execute arbitrary Python; it is being offered one explicitly defined capability with a known input structure.
That distinction matters when you move from a toy example to a real application. A narrowly defined tool such as lookup_customer or create_support_ticket gives you a much clearer security boundary than handing the agent unrestricted access to your application's internals.
Add streaming when the response should feel interactive
The basic example waits until the complete response is available. For an interactive application, that can make a long answer feel slower than it actually is. The SDK supports streaming events, including assistant message deltas, so your interface can display generated text as it arrives.
The session can be created with streaming enabled, then an event handler can print each piece of generated text:
session = await client.create_session(
on_permission_request=PermissionHandler.approve_all,
model="auto",
streaming=True,
tools=[calculator],
)
def on_event(event):
if event.type.value == "assistant.message_delta":
print(event.data.delta_content, end="", flush=True)
session.on(on_event)
await session.send_and_wait(
"Calculate 125 divided by 5 and explain the result."
)Streaming does not change what the agent can do. It changes how your application receives the response. For a terminal assistant, that means text appears progressively; in a web application, the same event-driven approach can be connected to whatever mechanism your frontend uses to display incremental output.
Turn the example into an application-specific agent
The calculator is useful because it isolates the tool mechanism, but real applications normally have more meaningful functions. Imagine a support dashboard where an agent needs to inspect an order and then create a support ticket. You could expose those operations as two separate tools, each with a small schema and a controlled implementation.
The model would handle the conversational part while your application remains responsible for the actual operations. A request such as "Check order 4812 and open a ticket if it has not shipped" could therefore become a sequence in which the agent calls an order lookup tool, examines the returned data, and calls a ticket tool only if the conditions require it. That is the useful boundary: the model reasons about which available capability is appropriate, while your application owns the implementation and permissions.
This pattern also fits naturally with the broader reliable agents problem. Reliability does not come from telling the model to be careful; it comes from constraining what the model can call, validating inputs, handling failures, and making consequential actions explicit.
Keep permissions narrow when you move beyond a demo
A calculator has almost no meaningful side effects. Production tools can have very different consequences. A function that changes a database record, sends an email, deletes a file, or executes a shell command should not automatically receive the same approval treatment as a read-only lookup.
The Copilot SDK has permission handling specifically because agent actions can require authorization. Its documentation also provides deployment patterns for applications serving multiple users, including per-user sessions and application-managed GitHub tokens.
One practical rule is to separate read and write capabilities. A tool that retrieves an invoice can usually be treated differently from one that changes its status. If a write operation is unavoidable, validate its parameters inside your application rather than trusting the model's interpretation of the user's request.
Persist sessions when the agent needs continuity
A short-lived command-line experiment can create a session, perform one task, and exit. A real application may need the agent to continue work after a restart or reconnect to an earlier conversation. The SDK supports persistent sessions when you provide your own session identifier, allowing state to be resumed later. GitHub's documentation describes this as a way to preserve conversation history, tool state, and planning context across restarts or other runtime changes.
For example, a session can be created with an application-defined identifier:
session = await client.create_session(
on_permission_request=PermissionHandler.approve_all,
model="auto",
session_id="customer-42-support",
)The important detail is that the identifier is yours to manage. If your application creates a random session identifier every time the user returns, there is no stable reference with which to resume the previous session. In a multi-user service, use an identifier that safely maps to your own account or conversation records rather than exposing sensitive information directly in the session name. ([GitHub][4])
Where the SDK fits when your agent grows
Once the basic client, session, and tool pattern works, the SDK provides several directions for expansion. Its current documentation covers hooks, custom agents, Model Context Protocol servers, reusable skills, cloud sessions, observability, authentication, and multi-tenant deployment patterns.
That makes it possible to grow the calculator example without throwing away the original architecture. A local Python function can become a database-backed service, a single session can become a per-user workflow, and a simple terminal response can become a streamed interface. The same underlying idea remains: define the capabilities your application owns, expose only the ones the agent needs, and let the Copilot runtime manage the agent loop around them.
There is also a useful architectural reason to pay attention to the SDK now. GitHub recently described a rewrite of the Copilot agent runtime into more than 800,000 lines of production Rust, replacing an earlier TypeScript and Node.js implementation and allowing the runtime to be embedded more directly across its SDK ecosystem. GitHub says the change was driven by startup, memory, throughput, embedding, and reliability requirements; it is not a blanket recommendation to rewrite TypeScript applications in Rust.
For a Python developer, the practical takeaway is simpler than the runtime's implementation story: you do not need to build that agent infrastructure yourself to experiment with an application-level Copilot agent. Start with one session and one narrowly scoped tool, verify the permission boundary, and add capabilities only when the application has a concrete need for them.
Written by


