Tool search that cuts input tokens by more than 97%, native voice agents, crash-resilient hosted agents, and a loop that turns production traces into better agents
Give an agent 1,000 tools and every turn starts with the model reading definitions it will mostly never call. Microsoft measured the fix: tool search in Microsoft Foundry Toolboxes cut input tokens by more than 97% at that catalog size, and by more than 60% with only 50 tools.

Tool search is one of more than a dozen changes in Foundry's September 2026 release, and most of them target problems that appear after an agent ships: tool catalogs grow, models change, users call in by phone, long jobs crash halfway, and quality drifts. Below, each component gets a diagram, the configuration that matters, its preview status, and the limits the documentation spreads across a few dozen pages.
The map: how the pieces fit together
Foundry Agent Service now has three agent types. Prompt agents are declarative: instructions, a model, and tools. Voice agents own a spoken conversation end to end. Hosted agents run your own container, and Foundry provides the endpoint, the scaling, and a Microsoft Entra identity for each agent. Channels and triggers call into that runtime, agents reach models, one toolbox endpoint, and state, and their traces feed an improvement loop while governance applies at every layer.

Several features in this release depend on hosted agents, so one more detail is worth having in mind. Each session runs in its own VM-isolated sandbox with a persistent $HOME and /files area. You own the code inside the sandbox; the platform owns the endpoint, identity, scaling, and session state around it.
Status matters as much as capability, because preview features come without an SLA:
- Generally available: tool search, the A2A tool, Routines, and Toolboxes for hosted agents.
- Available now: the GPT-6 family and Claude Opus 5.5 in Foundry, Microsoft Entra and Agent 365 governance actions enforced at runtime, the Microsoft Agent Framework updates, and the Foundry dev pack.
- General availability announced for late September 2026: the rubric evaluator, trace and synthetic dataset generation, and the agent optimizer.
- Public preview: voice agents, long-running resilience, Insights, network egress controls, the reminder tool, and Toolboxes for prompt agents.
- Coming: voice support in Foundry Toolkit for Visual Studio Code, and the Azure API Management AI Gateway tier integration (preview, October 2026).
1. Toolboxes: one governed endpoint for every tool
Most turns need a small slice of an agent's tools. A toolbox bundles tools and skills into one versioned resource that agents reach through a single managed MCP endpoint. It can hold MCP servers, OpenAPI tools, A2A connections, Azure AI Search, Web Search, Code Interpreter, File Search, Fabric IQ, Work IQ, skills, and tool search. Names are namespaced, so an MCP tool appears as {server_label}.{tool_name}.
Versioning is the part that changes how you operate tools. Every change creates an immutable version, and the toolbox keeps a default_version pointer. Agents call the consumer endpoint (/toolboxes/{name}/mcp), which always serves the default. You test a new version on its own developer endpoint (/toolboxes/{name}/versions/{v}/mcp), then promote or roll back by moving the pointer, with no agent redeploy.

from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import MCPToolboxTool, ToolSearchToolboxTool, WebSearchToolboxTool
project = AIProjectClient(
endpoint="https://<account>.services.ai.azure.com/api/projects/<project>",
credential=DefaultAzureCredential(),
)
toolbox_version = project.toolboxes.create_version(
name="service-desk-tools",
description="Web search, ITSM MCP server, and tool search",
tools=[
WebSearchToolboxTool(),
MCPToolboxTool(
server_label="itsm",
server_url="https://itsm-mcp.contoso.example",
require_approval="always",
project_connection_id="itsm-connection",
),
ToolSearchToolboxTool(),
],
)
print(toolbox_version.name, toolbox_version.version)A version can also carry a guardrail policy (policies.rai_config.rai_policy_name) that filters tool inputs and outputs at the tool layer, and skills attach as MCP Resources at skill://{name}. Three details save debugging time. The toolbox reports require_approval in _meta.tool_configuration but doesn't block the call, so your runtime has to pause and ask the user.
The platform may overwrite environment variables that start with FOUNDRY_. And toolbox endpoints only accept streaming tools/call requests.
Docs: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/toolbox?WT.mc_id=AZ-MVP-5000671
2. How tool search turns 1,000 tools into two
Adding {"type": "toolbox_search"} to a toolbox version changes what the model sees first. The initial tools/list hides every toolbox tool and exposes two meta-tools instead. tool_search(query, limit) takes a plain-language description of the capability the model needs and returns matching definitions (five by default, ten at most). call_tool(name, arguments) invokes a discovered tool by name. Pinned tools appear alongside them from the start.

The second meta-tool exists for a practical reason: many runtimes refuse to call a tool that wasn't in the original list, so call_tool gives them a registered, policy-aware way to dispatch. The model can search several times in one turn, and anything it discovers stays callable until the turn ends.
Ranking is where tool search becomes a search problem. Microsoft Learn describes BM25 over each tool's name, description, and parameters, and Microsoft's engineering write-up describes an enhanced sparse-retrieval pipeline that indexes argument names and descriptions up to three levels deep. On ToolRet (more than 44,000 tools and 7,000 queries), tool search matched a GPU-based reranker on Recall@10 for web and code queries (45.99% against 45.94%, and 39.56% against 38.23%), trailed it on customized queries (41.36% against 49.43%), and beat plain BM25s in all three.
The same evaluation found that misses were mostly editorial. Tools with generic descriptions ("get", "create", "REST API"), or with implementation vocabulary instead of the words users type, ranked poorly. Three settings address that:
additional_search_textadds search-only keywords the model never sees, so the token savings stay intact. Microsoft reports that tuned metadata improved retrieval hit rate by about 56% and end-to-end accuracy by about 55%, to within roughly 4% of the full-catalog baseline.pinkeeps a tool in everytools/list, and"*"pins a whole entry.- Auto-pinning surfaces each user's most-called tools after a short warmup.
{
"tools": [
{ "type": "toolbox_search" },
{
"type": "mcp",
"server_label": "analytics",
"server_url": "https://db-mcp.internal/sse",
"tool_configs": {
"execute_query": {
"pin": true,
"additional_search_text": "SQL analytics reporting dashboard warehouse queries"
},
"list_tables": {
"additional_search_text": "schema columns metadata table structure"
}
}
}
]
}Tool search pays off once a toolbox passes 10 to 15 tools, or when one agent serves many workflows. Either way, tell the model in its instructions to call tool_search before it concludes that a capability doesn't exist.
3. A2A: agents calling agents over an open protocol
The A2A tool, now generally available, lets a Foundry agent call other agents, including other Foundry agents, through the open Agent2Agent protocol instead of custom point-to-point integrations. Use the a2a tool for protocol version 1.0 (JSON-RPC); a2a_preview remains for version 0.3 integrations.

Credentials live in a project connection, with options that range from custom keys and OAuth2 to user Entra token passthrough, project managed identity, and agentic identity. To call another Foundry agent, point the connection at its /agents/{agent}/endpoint/protocols/a2a path with the audience https://ai.azure.com. The target needs incoming A2A enabled, and the caller needs the Foundry Agent Consumer role. Design around two limits: Foundry targets exchange text only, without streaming, and tasks are kept for 60 days after their last write.
Agents that run outside Foundry can be registered in Foundry Control Plane, which gives them a proxy URL and AI gateway monitoring. If you used Connected Agents in the classic Agents API, A2A and Foundry workflows replace it.
4. Routines: agents that start on their own
Routines, now generally available, start an agent from a one-time timer, a cron schedule, or an event, without a scheduler, queue, or webhook of your own. The first event sources are GitHub issue events and new Microsoft Teams channel messages. The action calls the agent through the Responses API, and you choose the identity: the routine creator's delegated access (for tools that use the user's OAuth identity) or the agent's own Entra identity.

# project = AIProjectClient(...) as in the toolbox example
routine = project.beta.routines.create_or_update(
routine_name="daily-summary",
description="Weekday morning summary",
enabled=True,
triggers={
"weekday-morning": {"type": "schedule", "cron_expression": "0 7 * * 1-5", "time_zone": "UTC"}
},
action={
"type": "invoke_agent_responses_api",
"agent_name": "service-desk",
"input": "Summarize the last 24 hours of tickets.",
},
)For work that outlasts one run, hosted agents get the reminder tool (preview). The agent schedules its own follow-up, and Foundry invokes the same agent on the same conversation after the delay.
Read more about Foundry Routines: https://medium.com/@daverendon/microsoft-foundry-routines-ai-agents-schedule-events-15e9a4b0c25e
Docs: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/use-routines?WT.mc_id=AZ-MVP-5000671
5. How Foundry voice agents handle a spoken turn
Most voice stacks attach speech to a text agent: speech-to-text, the agent, text-to-speech, a streaming layer, and separate monitoring for each. A voice agent (kind: voice, public preview) owns the whole spoken turn, and you build, deploy, trace, and evaluate it like any other Foundry agent. It deploys to web clients, Microsoft Teams, Teams Phone, and Twilio telephony, across 80+ languages and 140+ locales.

The model decides the pipeline. model_type is managed (for example gpt-realtime-2.1) or self_deployed, and the service derives the architecture from your choice. Speech-to-speech models take audio in and produce audio out in one model; the portal lists approximate latencies from about 400 ms for gpt-realtime-mini to about 800 ms for gpt-realtime-2.1. Cascaded text models such as gpt-5.4 transcribe, reason, and synthesize in about 1.5 seconds, in exchange for direct control over transcription and voice. Treat those figures as comparisons rather than guarantees, since tools, the network, and avatars all add time.
Turn detection is where many voice agents start to feel robotic. server_vad detects silence, semantic_vad uses OpenAI's end-of-turn model, and the Azure semantic variants include a multilingual option for callers who switch languages mid-call. Interruptions are on by default. For contact centers, use azure_deep_noise_suppression; for product names, add a phrase_list.
Two settings deserve an explicit decision. interim_response fills silence while a tool runs, because a quiet line sounds like a dropped call. And store is false by default; when it's true, the transcript, event timeline, and raw audio persist together, with no separate switch for audio.
# Requires azure-ai-projects 2.7.0 or later and AIProjectClient(..., allow_preview=True)
from datetime import timedelta
from azure.ai.projects.models import (
RealtimeAudioFormatsAudioPcm, VoiceAgentAudioConfig, VoiceAgentAudioInputConfig,
VoiceAgentDefinition, VoiceAgentServerVadTurnDetection,
VoiceAgentStaticInterimResponseConfig, VoiceModelType,
)
definition = VoiceAgentDefinition(
model_type=VoiceModelType.MANAGED,
model="gpt-realtime-2.1",
instructions="You are a service desk assistant. Use short sentences and ask one question at a time.",
)
definition.audio = VoiceAgentAudioConfig(
input=VoiceAgentAudioInputConfig(
format=RealtimeAudioFormatsAudioPcm(rate=24000),
turn_detection=VoiceAgentServerVadTurnDetection(
threshold=0.5, prefix_padding_ms=300, silence_duration_ms=500
),
)
)
definition.interim_response = VoiceAgentStaticInterimResponseConfig(
triggers=["latency", "tool"],
texts=["Let me check that for you.", "One moment while I look that up."],
latency_threshold_ms=timedelta(milliseconds=1500),
)
# project.agents.create_version(agent_name="service-desk-voice", definition=definition)Clients connect over wss://{project-endpoint}/agents/{agent-name}/endpoint/protocols/voice?api-version=v1 with a Microsoft Entra token. If you need full control, hosted agents also support real-time voice through the invocations_ws WebSocket protocol, with Pipecat, LiveKit, or Voice Live running inside your container.
6. What long-running resilience recovers, and what it doesn't
Research jobs, multi-tool reports, and waits for human approval all outlive a single HTTP request. Long-running resilience (public preview) keeps hosted-agent work going after the client disconnects and recovers it after the hosting process dies. The documentation separates three things that are easy to blur. Background execution lets work continue without an open connection. Resilient execution recovers work after the process is lost. Stream replay lets a reconnecting client pick up from a cursor.
Recovery is lease-based. Each unit of work carries a work identity (the job or conversation) and an input identity (one turn). The runtime persists the input, takes a lease on the work record, and renews the lease while your handler runs. If the process stops, a new process reclaims the lease and calls your handler again with the same identities and input.

That last sentence is the one to design for: recovery reenters the handler from the beginning. Local variables and the call stack are gone, and any side effect that already happened can happen again unless you guard it. Keep a small checkpoint index in task metadata (a framework checkpoint ID, the last completed phase, an idempotency key), keep large state in the hosted agent's durable state store or your own storage, and wrap anything that can't be repeated in a watermark:
# Illustrative pseudocode for a recovery-aware handler. Not an SDK API.
async def handle(work_id, input_id, payload, meta, store):
ckpt = await store.load(meta.get("checkpoint_ref")) # resume from durable progress
if ckpt is None or ckpt.phase < 1:
findings = await run_research(payload)
ckpt = await store.save(work_id, phase=1, findings=findings)
meta["checkpoint_ref"] = ckpt.ref
if meta.get("report_watermark") != "committed": # guard the non-repeatable step
meta["report_watermark"] = "pending"
await send_report(ckpt.findings, idempotency_key=f"{work_id}:{input_id}")
meta["report_watermark"] = "committed"
return summarize(ckpt.findings)Two more rules come straight from the documentation. Give each turn its own stream identity. And after a recovery, treat a later response.in_progress event as a snapshot reset in the client. On the Responses protocol, full crash recovery applies to stored background responses when you opt in.
Microsoft Agent Framework covers the framework side. Its workflow checkpointing resumes an interrupted background workflow from its latest durable checkpoint while Foundry preserves the hosted response.
This release also adds Agent Channel for messaging, AG-UI support, episodic procedural memory, and CodeAct with Hyperlight isolation, where the model writes one small program per turn instead of making a chain of tool calls. In a benchmark sample published with the framework, CodeAct finished 52.4% faster and used 63.9% fewer tokens than traditional tool wiring. That's one workload, so measure your own. The new Foundry dev pack installs the toolchain in one step.
👉Learn more about Foundry Dev Pack here: https://medium.com/towards-artificial-intelligence/microsoft-foundry-dev-pack-hosted-agent-d99faf25954e
On cost, hosted agents scale per session. The idle timeout runs from 2 to 60 minutes (15 by default), sandboxes range from 0.5 vCPU with 1 GiB to 2 vCPU with 4 GiB, and billing covers CPU and memory across active sessions, so an oversized version costs extra for every concurrent session.
7. The improvement loop: from production traces to a better agent
Shipping an agent is where the real work starts. Foundry now connects steps that used to be a pile of dashboards and ad hoc tests into one loop that starts and ends in production: observe, understand, define good, build tests, optimize, and validate.

Understand with Insights (public preview). Predefined evaluations catch the problems you already know about. Insights reads production traces to surface recurring issues, including ones you didn't anticipate, and attaches the supporting traces, a likely cause, and a suggested next step. Microsoft's example is an order-support agent that tells customers an order was delivered even when the status lookup failed. In one trace that looks like a fluke; across hundreds, it's a pattern.
Define "good" with the rubric evaluator. A rubric is a set of weighted criteria that an LLM judge scores from 1 to 5 for each response. The overall score is normalized to a 0 to 1 range, with a default pass threshold of 0.5. You can generate a rubric from an agent, a system prompt, or reference files, and ground it in production traces. The documentation recommends GPT-5 family judges, with gpt-5.4-mini as the best balance of quality and cost, and pairing rubrics with the built-in safety and groundedness evaluators.
Build tests from real traffic. Trace-based dataset generation uses intelligent sampling: it filters out low-intent traffic, uses MinHash to pick a diverse sample instead of near-duplicates, and handles personal data. Synthetic generation covers scenarios that production hasn't reached yet. Pin agent_version so a dataset doesn't mix behavior from several versions, and check generated_samples after each job, because max_samples (15 to 1,000) is only a ceiling.
Optimize and validate. The agent optimizer scores a baseline on your dataset, generates candidate configurations, scores them on the same dataset, and ranks them by a composite score from 0 to 1. It can rewrite instructions, refine skills in hosted agents, improve function-tool descriptions without touching types or required fields, and compare models.
That last target is how this release's model news becomes a decision rather than a debate: add the GPT-6 family or Claude Opus 5.5 to the candidate list (model_search_space in eval.yaml for hosted agents, or the wizard for prompt agents) and let your own data choose.

Microsoft's guidance for reading the result is worth keeping close. A gain under 0.03 is noise, 0.03 to 0.10 is moderate and worth deploying, 0.10 to 0.20 is significant, and anything above 0.20 usually means the baseline was weak.
Two cautions apply. Optimized instructions tend to be longer, which raises token cost. And every run executes your agent's tools for real, so point it at mocks or test endpoints.
Read more about the AI Agent ROI Framework here: https://blog.azinsider.net/ai-agent-roi-business-case-npv-azure-9e48f518a21b?sk=bba9ba1bab262cf5a9d0b242fddc76c8
8. Governance: treating agents as enterprise assets
Directory actions now reach the runtime. When you block, disable, delete, restore, or reassign the owner of an agent in Microsoft Entra or Agent 365, the Foundry runtime enforces it instead of only recording it in a directory.
Network egress controls (public preview). An agent's tool list says what it intends to call; an egress policy decides what its hosted sandbox can actually reach. Rules run top to bottom and the first match wins, with actions to allow, deny, transform headers, or rewrite the destination. A default action of Deny turns the policy into an allowlist, evaluation fails closed, and the platform's own domains are allowed automatically.

Start in Audit mode, which logs would-deny decisions to Application Insights while requests still go through, and switch to Enforced once the log is clean. Rules live in the agent's RAI policy and apply to hosted agents only. Allowing a host also grants no permission at that host, so keep API authorization separate from network policy.
Read more about Foundry egress controls here: https://medium.com/codetodeploy/microsoft-foundry-egress-controls-hosted-ai-agents-7944304236ee
AI Gateway for model access (preview, October 2026). A new AI Gateway tier in Azure API Management lets a platform team govern model access centrally, with quotas, rate limits, content safety, and logging. It publishes approved models into Foundry projects as Admin Connected Models, which developers then use for discovery, experimentation, agent development, and evaluation. Existing integrations with traditional API Management tiers keep working.
Proving a control works. The open-source run-assert-eval Skill chains three Microsoft projects inside your coding agent. Clarity (optional) helps surface failure modes. ASSERT turns a requirement into generated test cases, runs them against the live agent with OpenTelemetry traces, and judges the results. The Agent Control Specification (ACS) then adds a runtime control at defined points in the agent loop, and the Skill re-runs the same cases to show whether the fix held.

The idea worth borrowing is to measure two numbers separately: impermissible violations (harm) and permissible violations (legitimate work the agent refused or broke). Microsoft's published banking demo, which uses synthetic cases, shows why.
An ACS Rego policy cut impermissible authorization violations from 8% to 0%, while a defensive prompt only reached 6%. For coercion attempts that no structured field could detect, an ACS classifier and a hardened prompt both reached 0%, but the classifier kept permissible violations at 27% against 47% for the prompt. A prompt change is often the wrong control, and measuring both axes is how you find out.
A practical build order
Taken together, these features suggest a sequence for a new customer-operations agent. Treat it as a starting point:
- Start with a prompt agent, and shortlist two frontier models and one small model.
- Put every tool behind one toolbox with tool search on. Pin the few tools used on almost every turn, and enforce
require_approvalin your runtime. - Add a voice agent on the same toolbox, with interim responses and an explicit decision about
store. - Let other teams own their agents and call them over A2A, designed for text-only exchanges.
- Move long work into a hosted agent with resilient background responses, checkpoints in the state store, and a routine to start it.
- Attach an egress policy in Audit mode, review the decisions, then enforce it. Add an ASSERT gate to CI for your top risks.
- Run the loop on a schedule: Insights, the rubric, a fresh dataset from the pinned version, and the optimizer against mocked tools. Promote only when the gain clears 0.03 at a token cost you accept.
Sharp edges worth knowing before production
- Preview features (voice agents, resilience, Insights, egress controls, prompt-agent toolboxes) have no SLA.
- Toolboxes report
require_approval; your code has to enforce it. - Tool descriptions are now ranking features, so review them like search metadata.
- The optimizer calls your real tools during evaluation.
- For voice,
storepersists the transcript, timeline, and raw audio as one switch. - Hosted-agent sandbox size multiplies by concurrent sessions.
- A2A to Foundry targets is text-only and non-streaming today.
Final Thoughts
The pattern across this release is consistent: Foundry is moving the work around the agent into the platform. Tool catalogs, schedules, phone lines, crash recovery, evaluation data, and network policy used to be code every team wrote for itself. Now they're versioned resources you configure, observe, and roll back.
That changes where the effort goes. Calling a new model is rarely the hard part anymore. Knowing whether a change made your agent better is. So the one recommendation worth taking from all of this is to wire up tracing and write a rubric before you pick a model. With those two in place, every new model becomes a candidate you can score against your own traffic, and every improvement arrives with evidence.
If you try one thing this week, turn on tool search for your largest toolbox and compare the input tokens in your traces before and after. It's the quickest way to see where the platform is heading, in your own numbers.
References
Sources
- GPT-6 in Microsoft Foundry: https://azure.microsoft.com/en-us/blog/gpt-6-astra-sol-and-luna-for-production-agents-in-microsoft-foundry/
- Tool search evaluation: https://commandline.microsoft.com/tool-search-toolboxes-foundry/
- Routines announcement: https://devblogs.microsoft.com/foundry/from-chatbots-to-automated-assistants-routines-in-microsoft-foundry-are-now-generally-available/
- Insights announcement: https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/insights-in-foundry-turns-agent-traces-into-action/4559634
- ASSERT and ACS write-up: https://commandline.microsoft.com/safety-requirements-failure-paths-assert-acs/
- Agent Framework CodeAct sample and results: https://devblogs.microsoft.com/agent-framework/microsoft-agent-framework-at-build-2026-announce/
Microsoft Learn: runtime and tools
- Foundry Agent Service overview: https://learn.microsoft.com/en-us/azure/foundry/agents/overview?WT.mc_id=AZ-MVP-5000671
- Hosted agents: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/hosted-agents?WT.mc_id=AZ-MVP-5000671
- Long-running resilience: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/long-running-agent-resilience?WT.mc_id=AZ-MVP-5000671
- Durable state store: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-state-store?WT.mc_id=AZ-MVP-5000671
- Create and manage a toolbox: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/toolbox?WT.mc_id=AZ-MVP-5000671
- Enable tool search: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/tool-search?WT.mc_id=AZ-MVP-5000671
- A2A tool: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/agent-to-agent?WT.mc_id=AZ-MVP-5000671
- Enable incoming A2A: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/enable-agent-to-agent-endpoint?WT.mc_id=AZ-MVP-5000671
- Use routines: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/use-routines?WT.mc_id=AZ-MVP-5000671
- Reminder tool: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/reminder-tool?WT.mc_id=AZ-MVP-5000671
Microsoft Learn: voice
- Configure a voice agent: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/configure-voice-agent?WT.mc_id=AZ-MVP-5000671
- Voice agent quickstart: https://learn.microsoft.com/en-us/azure/foundry/agents/quickstarts/prompt-voice-agent?WT.mc_id=AZ-MVP-5000671
- Build a voice agent with hosted agents: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/build-voice-agent?WT.mc_id=AZ-MVP-5000671
- Voice agent observability: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/voice-agent-observability?WT.mc_id=AZ-MVP-5000671
Microsoft Learn: improvement loop
- Agent tracing: https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept?WT.mc_id=AZ-MVP-5000671
- Rubric evaluators: https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/rubric-evaluators?WT.mc_id=AZ-MVP-5000671
- Convert traces into evaluation datasets: https://learn.microsoft.com/en-us/azure/foundry/observability/how-to/traces-to-dataset?WT.mc_id=AZ-MVP-5000671
- Agent optimizer overview: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-optimizer-overview?WT.mc_id=AZ-MVP-5000671
- Agent optimizer costs: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-optimizer-costs?WT.mc_id=AZ-MVP-5000671
Microsoft Learn: governance
- Hosted agent guardrails and egress controls: https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/add-hosted-agent-guardrails?WT.mc_id=AZ-MVP-5000671
- Agent identity: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/agent-identity?WT.mc_id=AZ-MVP-5000671
- AI gateway in Azure API Management: https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities?WT.mc_id=AZ-MVP-5000671
GitHub
- Foundry samples: https://github.com/microsoft-foundry/foundry-samples
- Microsoft Agent Framework: https://github.com/microsoft/agent-framework
- Clarity: https://github.com/microsoft/clarity-agent
- ASSERT: https://github.com/responsibleai/ASSERT
- Agent Control Specification: https://github.com/microsoft/agent-governance-toolkit/tree/main/policy-engine
Research
- ToolRet benchmark: https://arxiv.org/abs/2503.01763