Skip to content

The Edge of the Cyber World See the latest

Apps

Microsoft Agent Framework Setup: 13 Steps, 90 Min [2026]

Microsoft folded Semantic Kernel and AutoGen into a single production SDK called Microsoft Agent Framework, and as of September 2026 it has settled into a stable, versioned release that developers can actually build on. The Python package sits at version 1.19.0, the .NET package at 1.22.0, both shipped September 18, 2026, and the framework crossed its 1.0 milestone back in April. If you tried an early preview build and walked away confused by shifting class names, this guide walks through the current, stable API from a clean install to a working multi-agent project. Expect real code you can run today, not conceptual diagrams.

This tutorial covers 13 concrete steps: installing the package, wiring up an OpenAI-backed agent, adding tools, memory, and streaming, connecting to Azure and Microsoft Foundry, orchestrating multiple agents through workflows, and packaging the result for production. Total hands-on time runs about 90 minutes if you already have Python installed and an OpenAI API key in hand.

What Is Microsoft Agent Framework?

Microsoft Agent Framework (often shortened to Agent Framework or MAF) is an open source SDK for building AI agents and multi-agent workflows in Python, .NET, and, as of a public preview, Go. Microsoft describes it as the direct successor to both Semantic Kernel and AutoGen, built by the same internal teams. The pitch is straightforward: AutoGen contributed simple, fast-to-write agent abstractions and multi-agent conversation patterns, while Semantic Kernel contributed enterprise features like session-based state management, type safety, middleware, and telemetry. Agent Framework merges both, then adds graph-based workflows for explicit control over how agents and functions execute in sequence, as laid out in the framework’s official overview documentation.

The project is MIT-licensed and hosted at github.com/microsoft/agent-framework, where it has accumulated roughly 13,800 stars. That is a fraction of LangChain’s install base, but Agent Framework is not really competing on general-purpose popularity. It is aimed squarely at teams already inside the Microsoft and Azure ecosystem who want first-party support for Azure OpenAI, Microsoft Foundry, and enterprise identity, without giving up compatibility with plain OpenAI, Anthropic, Amazon Bedrock, Google Gemini, and local models through Ollama.

Microsoft announced the 1.0 milestone in April 2026, marking the point where the API surface stopped shifting week to week and became something teams could depend on for a real product. Before that, developers building on early previews had to track breaking changes across near-weekly releases, which is part of why so many older blog posts and tutorials floating around the web reference class names and import paths that no longer exist. If you land on a guide that imports from autogen_agentchat or references a standalone semantic_kernel agent class instead of agent_framework, you are reading documentation for a predecessor project, not the current framework.

Since the 1.0 release, updates have followed a steady monthly cadence rather than the breaking-change churn of the preview period. September 2026 alone shipped three Python point releases, 1.17.0, 1.18.0, and 1.19.0, alongside a matching .NET 1.22.0 build, all adding incremental provider support and bug fixes without disturbing the core agent and workflow APIs this tutorial relies on.

Four pieces make up the framework. Agents are the individual LLM-backed workers that call tools and generate responses. A newer addition called Harness Agent bundles planning, todo tracking, context compaction, and file access for long multi-step tasks. Workflows connect agents and plain functions through explicit graphs rather than leaving every decision to the model. Integrations tie all of that to model providers, hosted tools, evaluation services, and UI frameworks. You will touch the first three directly in this tutorial.

Model Providers and Pricing: What You Can Connect To

One reason teams evaluate Agent Framework alongside LangChain or CrewAI is that it does not lock you into a single model vendor despite the Microsoft branding. The core framework ships free under the MIT license. What you pay for is whichever model provider you point an agent at, and that list is wider than the name suggests.

Provider Python package Typical use case
OpenAI agent-framework-openai Direct OpenAI API access, Responses or Chat Completions
Azure OpenAI agent-framework-openai (same client, different arguments) Enterprise deployments with Azure identity and billing
Microsoft Foundry agent-framework-foundry Service-managed agents, Foundry-hosted models
Anthropic Provider connector via Foundry or dedicated package Claude models for teams that want an alternative to GPT
Amazon Bedrock Provider connector AWS-native deployments that already run on Bedrock
Google Gemini Provider connector Teams standardized on Google Cloud AI services
Ollama Local, OpenAI-compatible client via base_url Fully local inference with no per-token cost

Switching providers rarely means rewriting your agent logic. Because OpenAIChatClient accepts a base_url parameter, it can point at any OpenAI-compatible endpoint, including a local Ollama server or LM Studio instance, without touching your agent, tool, or session code. That portability is worth testing early in a project, before you have built dozens of agents against one provider’s specific quirks.

Prerequisites: What You Need Before You Start

You do not need an Azure subscription to follow most of this guide. A plain OpenAI API key covers steps 1 through 8 and step 10. Azure OpenAI and Microsoft Foundry are optional add-ons covered in step 9. Here is what to have ready before you start the clock.

Requirement Minimum version Notes
Python 3.10 or newer Required by the agent-framework 1.19.0 wheel on PyPI
pip 23.x or newer Run python -m pip install --upgrade pip first if unsure
agent-framework (Python package) 1.19.0 Released September 18, 2026, core package plus optional extras
Microsoft.Agents.AI (.NET package) 1.22.0 Released September 18, 2026, for the C# path, see the NuGet package page
.NET SDK 8.0 or newer Only needed if you follow the C# examples instead of Python
OpenAI API key Active account with billing Create one at platform.openai.com, a paid tier avoids rate-limit walls
Azure CLI (optional) Latest Only needed for the Azure OpenAI or Microsoft Foundry steps
Code editor VS Code or similar Python and .NET extensions recommended but not required

Budget for real API spend once you move past the first hello-world call. A typical gpt-4o-mini or gpt-5.4-mini agent run in this tutorial costs a fraction of a cent per turn, but hosted tools like code interpreter and web search bill separately, and workflows that fan out to several agents multiply the token count fast. Keep an eye on your OpenAI usage dashboard while you test.

Microsoft Agent Framework vs LangChain, CrewAI, and LangGraph

Before committing to a framework, it helps to see where Agent Framework sits relative to the tools most Python developers already know. None of these projects are going away, and picking one rarely locks you out of the others permanently, but the starting philosophy differs enough to matter for a greenfield project.

Framework Approx. GitHub stars Primary strength Best fit
LangChain ~147,000 Broadest ecosystem of integrations and community examples Teams that want maximum provider and tool coverage out of the box
CrewAI ~56,500 Role-based crews with a simple mental model Task delegation between a small number of specialist agents
LangGraph ~38,700 Durable, stateful graph execution Long-running workflows that need checkpointing and replay
Microsoft Agent Framework ~13,800 Enterprise state management plus explicit workflow graphs Azure-centric teams migrating off Semantic Kernel or AutoGen

Star counts favor the older, broader projects simply because they launched earlier and cover more use cases. What the table does not show is production usage, and by that measure the gap narrows: LangGraph alone reportedly clears several million monthly PyPI downloads, evidence that GitHub stars and real deployment volume do not move in lockstep. Readers who want the deeper breakdown of LangChain, LlamaIndex, and CrewAI star trajectories can find it in tech-insider.org’s dedicated framework comparison. The short version for this tutorial: pick Agent Framework when Azure identity, Microsoft Foundry, or a Semantic Kernel migration is already on your roadmap, and treat the workflow engine as the differentiator worth testing first.

Step 1: Verify Python and Set Up a Virtual Environment

Confirm your Python version meets the 3.10 floor before installing anything. Mixing a system-wide install with a project-specific one is the single most common source of “module not found” errors later, so isolate the environment from the start.

python3 --version
# Should print Python 3.10.x or higher

python3 -m venv maf-tutorial
source maf-tutorial/bin/activate   # on Windows: maf-tutorialScriptsactivate

python -m pip install --upgrade pip

If python3 --version comes back below 3.10, install a current Python release from python.org or via your OS package manager before continuing. Agent Framework 1.19.0 simply will not resolve its dependencies on an older interpreter, and pip will fail with a version-conflict error rather than a friendly warning.

Step 2: Install the agent-framework Package

The core package installs the agent, thread, and workflow primitives. Provider-specific clients ship as separate optional packages so you only pull in what you need. For this tutorial you want the core package plus the OpenAI provider.

pip install agent-framework
pip install agent-framework-openai
pip install python-dotenv

# Pin versions for a reproducible tutorial environment
pip install agent-framework==1.19.0 agent-framework-openai

Run pip show agent-framework afterward to confirm the version matches what you expect, or check the current release directly on PyPI. If you plan to connect to Azure OpenAI or Microsoft Foundry later, also install azure-identity now so you are not stopping mid-tutorial to fetch one more dependency.

Step 3: Get an OpenAI API Key and Configure Environment Variables

Generate a key at platform.openai.com/api-keys and scope it to a specific project rather than reusing an organization-wide key. Agent Framework does not load .env files automatically, which trips up a lot of first-time users, so you either call load_dotenv() explicitly or export the variables in your shell.

# .env file in your project root
OPENAI_API_KEY="sk-your-key-here"
OPENAI_CHAT_MODEL="gpt-4o-mini"
# at the top of your Python script
from dotenv import load_dotenv
load_dotenv()

Pick a model name that matches your account tier. gpt-4o-mini is inexpensive and works for every example in this tutorial, but if your account has access to newer models like GPT-6 Sol or GPT-6 Luna, swap the value and everything downstream keeps working unchanged, since Agent Framework treats the model identifier as a plain string passed to the provider.

Step 4: Write Your First Agent

An agent in Agent Framework needs three things at minimum, according to Microsoft’s official get-started documentation: a chat client, a set of instructions, and something to run. The Python API centers on a class named Agent, paired with a provider-specific client. OpenAIChatClient targets OpenAI’s newer Responses API, which is the recommended default because it supports hosted tools like code interpreter and web search. Create a file called hello_agent.py.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    agent = Agent(
        client=OpenAIChatClient(),
        name="HelloAgent",
        instructions="You are a friendly assistant. Keep answers brief.",
    )
    result = await agent.run("What is the largest city in France?")
    print(result)

if __name__ == "__main__":
    asyncio.run(main())

Notice that OpenAIChatClient() takes no arguments here. It reads OPENAI_API_KEY and OPENAI_CHAT_MODEL straight from the environment. That convenience is also a common source of confusion, since passing an explicit model= argument silently overrides whatever the environment variable says.

Step 5: Run the Agent and Inspect the Output

Run the script from your activated virtual environment.

python hello_agent.py

Expect output close to this:

Paris is the largest city in France, with a population of roughly 2.1 million
within city limits and over 11 million in the greater metropolitan area.

The exact wording will vary since you are calling a live model, but a plain text answer confirms the full chain worked: environment variables loaded, the client authenticated against OpenAI, the agent sent your prompt, and the response printed. If you get an empty result or a stack trace instead, jump ahead to the troubleshooting section before continuing, since every later step builds on this one working correctly.

Step 6: Add a Custom Function Tool

Tools are what separate an agent from a chatbot. Agent Framework exposes a @tool decorator that turns any typed Python function into something the model can call. The framework reads the function’s type hints and docstring to build the tool schema automatically, so keep both accurate.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent, tool
from agent_framework.openai import OpenAIChatClient

load_dotenv()

@tool
def get_weather(location: str) -> str:
    """Get the current weather for a given city."""
    # In production, call a real weather API here
    return f"The weather in {location} is sunny, 25 degrees Celsius."

async def main():
    agent = Agent(
        client=OpenAIChatClient(),
        name="WeatherAgent",
        instructions="You are a weather assistant. Use the tool for weather questions.",
        tools=get_weather,
    )
    result = await agent.run("What's the weather in Tokyo?")
    print(result)

if __name__ == "__main__":
    asyncio.run(main())

Run it and you should see the agent invoke get_weather behind the scenes, then weave the return value into a natural-language answer instead of just echoing the raw string. You can pass a list of functions to tools= once you have more than one, and the model decides at runtime which ones, if any, a given prompt requires.

Step 7: Add Multi-Turn Memory With Sessions

Every call to agent.run() is stateless by default, which surprises people coming from chat-style APIs. To carry conversation history across turns, create a session and pass it into every subsequent call.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    agent = Agent(
        client=OpenAIChatClient(),
        name="MemoryAgent",
        instructions="You are a helpful assistant.",
    )
    session = agent.create_session()

    first = await agent.run("My name is Alice", session=session)
    print(first)

    second = await agent.run("What's my name?", session=session)
    print(second)  # correctly answers "Alice"

if __name__ == "__main__":
    asyncio.run(main())

Reuse the same session object for the length of one logical conversation, then discard it or persist it depending on your application. Sessions hold conversation state in memory by default. If your app needs durability across process restarts, serialize the session yourself or move to a service-managed option like Microsoft Foundry Agent Service, which stores history server-side instead of in your application’s memory.

A useful mental model: a session is a container for turns, not a database. Treat it like you would a request-scoped object in a web framework. Spin one up when a user starts interacting, keep passing it into every agent.run() call for that user’s conversation, and let it go out of scope when the conversation ends. For chat applications with thousands of concurrent users, store a lightweight session identifier per user in your own database and reconstruct or fetch the underlying state on demand rather than keeping every session object resident in memory indefinitely.

Step 8: Stream Responses in Real Time

For anything user-facing, streaming tokens as they arrive beats waiting for the full response. Pass stream=True to agent.run() and iterate over the result asynchronously.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    agent = Agent(
        client=OpenAIChatClient(),
        name="StoryAgent",
        instructions="You are a creative storyteller.",
    )
    print("Agent: ", end="", flush=True)
    async for chunk in agent.run("Tell me a short story about a robot.", stream=True):
        if chunk.text:
            print(chunk.text, end="", flush=True)
    print()

if __name__ == "__main__":
    asyncio.run(main())

Each chunk is a partial update object, and not every chunk carries text, since some represent tool-call events or metadata. Always guard with if chunk.text before printing, or you will end up with blank lines scattered through your output.

Step 9: Connect to Azure OpenAI or Microsoft Foundry

This step is optional and only relevant if your organization runs models through Azure. As of the current release, Azure OpenAI uses the same agent_framework.openai client classes as direct OpenAI. What changes is which arguments you pass: supply azure_endpoint, credential, and api_version to route the client to Azure instead of OpenAI’s public API.

pip install azure-identity
import asyncio, os
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient
from azure.identity import AzureCliCredential

async def main():
    agent = Agent(
        client=OpenAIChatClient(
            model=os.environ["AZURE_OPENAI_CHAT_MODEL"],
            azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"],
            api_version=os.getenv("AZURE_OPENAI_API_VERSION"),
            credential=AzureCliCredential(),
        ),
        name="AzureAgent",
        instructions="You are a helpful assistant.",
    )
    result = await agent.run("Hello!")
    print(result)

asyncio.run(main())

Run az login before executing this script, since AzureCliCredential depends on an authenticated Azure CLI session. For a fully managed alternative, Microsoft Foundry offers FoundryChatClient, which points at a Foundry project endpoint instead of a raw Azure OpenAI resource and can host agents server-side through the Foundry Agent Service rather than keeping all state in your own process.

Step 10: Add Hosted Tools for Web Search and Code Execution

Beyond your own custom functions, the Responses-based OpenAIChatClient exposes hosted tools that run on OpenAI’s infrastructure rather than yours: code interpreter, file search, web search, image generation, a hosted shell, and hosted MCP connections. Each one is created through a get_*_tool() method on the client.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    client = OpenAIChatClient()
    code_interpreter = client.get_code_interpreter_tool()
    web_search = client.get_web_search_tool()

    agent = Agent(
        client=client,
        name="PowerAgent",
        instructions="You can search the web and run code to answer questions.",
        tools=[code_interpreter, web_search],
    )
    result = await agent.run(
        "Search for the current Python latest stable version, then write code "
        "that prints it."
    )
    print(result)

asyncio.run(main())

Hosted tools only work with the Responses-based client. If you are using OpenAIChatCompletionClient for broader model compatibility, you lose access to code interpreter, file search, and hosted MCP, and keep only function tools and web search. Check the tool support matrix in OpenAI’s Agent Framework provider docs before committing to one client type for a project that needs specific hosted tools.

Step 11: Orchestrate Multiple Agents With Workflows

A single agent is fine for open-ended, conversational tasks. Once a process has well-defined steps, or you need multiple specialist agents to hand work back and forth, switch to a workflow. Workflows are graph-based, connecting agents and ordinary functions through explicit execution paths instead of leaving the routing decision entirely to a model.

Orchestration pattern How it behaves
Sequential Agents run in a fixed order, each consuming the previous stage’s output
Concurrent Independent agents run in parallel and their results are combined
Handoff One agent transfers control to a specialist mid-conversation
Group chat Multiple agents participate in a single, managed conversation
Magentic-One A multi-agent coordination strategy drawn from Microsoft Research and AutoGen

A minimal sequential setup looks like this: a research agent gathers facts, then hands its output to a writer agent that turns them into a summary.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    client = OpenAIChatClient()

    researcher = Agent(
        client=client,
        name="Researcher",
        instructions="List three current facts about the given topic, plainly.",
    )
    writer = Agent(
        client=client,
        name="Writer",
        instructions="Turn the given facts into a two-sentence summary.",
    )

    facts = await researcher.run("Microsoft Agent Framework")
    summary = await writer.run(f"Facts: {facts}")
    print(summary)

asyncio.run(main())

That example chains two agents by hand, which works for a quick prototype. For anything more complex, reach for the framework’s dedicated workflow builder instead of manually stitching calls together in your own code. A graph-based workflow describes the same researcher-then-writer pipeline as a set of explicit edges, which makes the execution path visible, inspectable, and reusable across different entry points.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent, WorkflowBuilder
from agent_framework.openai import OpenAIChatClient

load_dotenv()

async def main():
    client = OpenAIChatClient()

    researcher = Agent(
        client=client,
        name="Researcher",
        instructions="List three current facts about the given topic, plainly.",
    )
    writer = Agent(
        client=client,
        name="Writer",
        instructions="Turn the given facts into a two-sentence summary.",
    )

    workflow = (
        WorkflowBuilder()
        .add_edge(researcher, writer)
        .set_start(researcher)
        .build()
    )

    result = await workflow.run("Microsoft Agent Framework")
    print(result)

asyncio.run(main())

The exact builder API surface can shift slightly between minor releases, so check the workflows documentation for your installed version before copying this pattern into a large project. What stays constant across versions is the underlying idea: define nodes as agents or plain functions, connect them with edges, and let the workflow engine handle execution order, retries, and checkpointing instead of writing that orchestration logic yourself by hand.

Step 12: Add Middleware, Approval Gates, and Telemetry

Production agents need guardrails beyond “call the model and hope.” Agent Framework carries over Semantic Kernel’s middleware system, which lets you intercept agent actions before or after they execute. Common uses include logging every tool call, redacting sensitive fields, enforcing rate limits, and requiring human approval before a function tool with side effects actually runs.

Tool approval is handled by the framework’s function-invoking chat client, so it applies to any function tool call regardless of which underlying API you chose. For observability, the framework emits OpenTelemetry-compatible spans, and cache-usage metrics such as cache_read_input_token_count and cache_creation_input_token_count map directly onto standard gen_ai.usage.* attributes, so existing OpenTelemetry dashboards pick them up without custom instrumentation. Wire that into whatever tracing backend your team already runs before you ship an agent that touches production data, not after.

Step 13: Package the Project and Prepare for Production

Freeze your dependencies once the prototype works. Pin exact versions in a requirements.txt, never commit your .env file, and separate configuration from code so the same script runs against a development key and a production one without edits.

# requirements.txt
agent-framework==1.19.0
agent-framework-openai
python-dotenv
azure-identity

For a production deployment, run separate agent and session instances for each concurrent user rather than sharing one client across requests. A single OpenAIChatClient instance can safely serve concurrent async calls on the same event loop, including overlapping streaming and non-streaming requests, but that guarantee does not extend to OpenAIChatCompletionClient, and it never extends to mutating a client’s configuration while calls are in flight. Containerize the app, set your API keys through your platform’s secret manager instead of plain environment files, and add the middleware-based logging from step 12 before the first real user touches it.

Complete Working Project: A Research-and-Summarize Agent

Put the pieces together into one runnable file that combines a custom tool, a hosted web-search tool, and a session, producing a small but genuinely useful agent that looks up a topic and returns a structured brief.

import asyncio
from dotenv import load_dotenv
from agent_framework import Agent, tool
from agent_framework.openai import OpenAIChatClient

load_dotenv()

@tool
def save_note(topic: str, summary: str) -> str:
    """Save a research summary to a local notes file."""
    with open("research_notes.txt", "a", encoding="utf-8") as f:
        f.write(f"## {topic}n{summary}nn")
    return f"Saved notes on {topic}."

async def main():
    client = OpenAIChatClient()
    web_search = client.get_web_search_tool()

    agent = Agent(
        client=client,
        name="ResearchBriefAgent",
        instructions=(
            "You research a topic using web search, write a three-sentence "
            "brief, then save it using the save_note tool."
        ),
        tools=[web_search, save_note],
    )

    session = agent.create_session()
    topic = "Microsoft Agent Framework"
    result = await agent.run(
        f"Research the topic '{topic}' and save your brief.", session=session
    )
    print(result)

if __name__ == "__main__":
    asyncio.run(main())

Running this produces a printed confirmation plus a new research_notes.txt file in your working directory, containing a dated section for the topic you queried. Swap the hardcoded topic variable for user input, wrap the whole thing in a loop with a persistent session, and you have a minimal but functional research assistant you can extend with more tools, a second summarizing agent, or a workflow that runs on a schedule.

Common Pitfalls When Building With Microsoft Agent Framework

A handful of mistakes account for most of the friction developers report when adopting Agent Framework for the first time.

  • Assuming .env files load automatically. They do not. Forgetting load_dotenv() is the single most common cause of “missing API key” errors on a first run.
  • Mixing Responses and Chat Completions tool expectations. Code interpreter, file search, and hosted MCP only work with OpenAIChatClient. Developers who start with OpenAIChatCompletionClient for its broader model compatibility, then add a hosted tool call, hit a silent capability gap.
  • Treating agent.run() as stateful by default. Without an explicit session, every call is a fresh conversation. Multi-turn context loss is easy to misdiagnose as a model problem when it is actually a missing session object.
  • Assuming Semantic Kernel or AutoGen code drops in unchanged. Agent Framework is a genuine successor, not a compatibility shim. Migrating an existing app means adapting agent construction, tool registration, and orchestration code, not just swapping an import path.
  • Sharing one client across concurrent requests without isolation. OpenAIChatCompletionClient does not guarantee safe concurrent use the way OpenAIChatClient does. Create a separate agent and session per concurrent run in any multi-user deployment.
  • Installing the core package and expecting Foundry or Azure to work out of the box. Provider connectors for Azure OpenAI, Microsoft Foundry, and others ship as separate optional packages that need their own pip install.

Troubleshooting Common Errors

Work through these when something breaks, roughly in the order you are likely to hit them.

  • “OPENAI_API_KEY not set” or similar KeyError. Confirm load_dotenv() runs before you construct OpenAIChatClient(), and that your .env file sits in the directory you launched Python from.
  • ImportError on agent_framework.openai. You installed the core agent-framework package but skipped the provider extra. Run pip install agent-framework-openai.
  • ModuleNotFoundError for azure.identity. Only needed for Azure OpenAI or Foundry examples. Run pip install azure-identity and confirm az login succeeded first.
  • Agent forgets earlier turns in a conversation. You are almost certainly calling agent.run() without passing the same session object across calls.
  • Tool never gets called even though the model should need it. Check that your function has type hints on every parameter and a docstring. Agent Framework builds the tool schema from both, and a missing docstring can produce a schema the model does not understand.
  • Streaming loop prints blank lines. Not every chunk carries text. Always check if chunk.text before printing inside a streaming loop.
  • pip fails to resolve agent-framework dependencies. Almost always a Python version below 3.10. Run python3 --version and upgrade if needed.
  • Azure requests silently route to OpenAI instead of Azure. If OPENAI_API_KEY is set in your environment alongside Azure variables, the generic client stays on OpenAI unless you explicitly pass azure_endpoint or credential.
  • DefaultAzureCredential works locally but fails in production. It is meant for development. Switch to a specific credential type, such as a managed identity, before deploying.
  • Hosted tool calls fail with a capability error. You are likely using OpenAIChatCompletionClient, which does not support code interpreter, file search, or hosted MCP. Switch to OpenAIChatClient.

Advanced Tips for Production Agents

Once the basics work, a few refinements separate a demo from something you can actually run at scale. First, use explicit prompt-cache breakpoints on any stable block of context you repeat across requests, such as a long system prompt or a policy document. OpenAIChatClient supports prompt_cache_key and prompt_cache_options, and on supported models the savings on repeated input tokens are substantial enough to change your unit economics at volume.

Second, decide deliberately between an agent and a workflow rather than defaulting to whichever one you learned first. If you can write a plain function to handle a task deterministically, do that instead of routing it through an LLM call at all. Reserve agents for genuinely open-ended or conversational work, and reserve workflows for processes with well-defined steps where you need explicit control over execution order.

Third, if you are adapting a non-standard OpenAI-compatible endpoint, such as a self-hosted vLLM server that returns reasoning content in a nonstandard field, use the response_parser and message_preparer hooks on OpenAIChatCompletionClient rather than subclassing the client. Both hooks run after the framework’s default conversion, so you only need to handle the delta between your endpoint’s format and the standard one.

Finally, treat the migration guides seriously if you are moving off Semantic Kernel or AutoGen. Microsoft ships dedicated migration documentation for both paths precisely because the API surface changed enough that a mechanical find-and-replace will not get you a working app.

A fifth tip worth calling out separately: pin your dependency versions before you deploy anything that generates revenue or touches customer data. Agent Framework’s monthly release cadence is a good thing for staying current, but a minor version bump landing automatically in a CI pipeline can change default behavior in ways that only show up once real traffic hits the new code path. Pin agent-framework==1.19.0 and its provider extras exactly, test the upgrade in staging against your actual prompts and tools, and bump the pin deliberately rather than letting an unbounded install pull in whatever shipped that week.

Sixth, build a small evaluation harness before you scale past a handful of hand-tested prompts. Even a simple script that runs twenty representative user queries against your agent and checks the outputs against expected patterns will catch regressions that manual testing misses, especially after a model swap or a prompt tweak. Agent Framework does not ship an opinionated evaluation framework out of the box, so this is worth building early rather than bolting on after a production incident.

Frequently Asked Questions

Is Microsoft Agent Framework free to use?

The framework itself is open source under the MIT license and free to install from PyPI or NuGet. You still pay for whatever model provider you connect it to, whether that is OpenAI API usage, Azure OpenAI token costs, or Microsoft Foundry inference charges. Running models locally through Ollama avoids per-token fees but adds your own hardware and hosting costs instead.

Does Microsoft Agent Framework replace Semantic Kernel and AutoGen?

Yes. Microsoft describes it as the direct successor to both, built by the same teams, and positions it as the forward-looking choice for new projects. Existing Semantic Kernel or AutoGen applications do not become source-compatible automatically. Migrating requires adapting agent construction, model-client configuration, and orchestration code, which is why Microsoft publishes separate migration guides for each.

Which programming languages does Agent Framework support?

Python and .NET are the primary, fully supported languages. Go support exists as a public preview, with the caveat that declarative agents, retrieval-augmented generation, and functional workflows are not yet available on that path.

Can I use Microsoft Agent Framework without Azure?

Yes. The framework works directly with OpenAI, Anthropic, Amazon Bedrock, Google Gemini, and local models through Ollama, with no Azure subscription required. Azure OpenAI and Microsoft Foundry are supported integrations, not requirements.

What is the difference between an agent and a workflow?

An agent handles open-ended, conversational tasks where the model decides what to do next, including which tools to call. A workflow is a graph-based structure that connects agents and plain functions through explicit, predefined execution paths, which suits processes with well-defined steps where you need deterministic control over ordering and routing.

Why does my agent forget what I said earlier in the conversation?

Every call to agent.run() is stateless unless you pass a shared session object created with agent.create_session(). Reuse that same session across every turn of a given conversation to retain context.

Should I use OpenAIChatClient or OpenAIChatCompletionClient?

OpenAIChatClient targets the newer Responses API and is Microsoft’s recommended default, since it unlocks hosted tools like code interpreter, file search, web search, image generation, and hosted MCP. Reach for OpenAIChatCompletionClient only when you need broad compatibility with OpenAI-compatible endpoints that do not support Responses, or when migrating an existing Chat Completions integration.

How does Microsoft Agent Framework compare to LangChain for a new project?

LangChain has a much larger ecosystem of integrations and community examples, with roughly 147,000 GitHub stars versus Agent Framework’s roughly 13,800. Agent Framework’s advantage shows up in enterprise state management, explicit workflow graphs, and first-party Azure and Microsoft Foundry support. Teams already committed to Azure or migrating off Semantic Kernel tend to get more value from Agent Framework. Teams that want the widest possible provider and tool coverage often lean toward LangChain instead.

Can I run Microsoft Agent Framework agents fully offline?

Yes, with a local model server. Point OpenAIChatClient at an Ollama or LM Studio endpoint through the base_url parameter and the same agent, tool, and session code runs against a locally hosted model instead of a hosted API. You lose access to OpenAI’s hosted tools like code interpreter and hosted MCP, since those run on OpenAI’s infrastructure, but function tools, sessions, and streaming all continue to work.

What happened to the ChatAgent class from earlier previews?

Current Microsoft documentation and the stable 1.x package line center on a class named Agent rather than ChatAgent. If you find older tutorials, blog posts, or cached AI-generated answers referencing ChatAgent, treat them as describing a preview-era API and cross-check the current class name against the official get-started guide for your installed version before copying code.

Related Coverage

Source: Tech Insider