# Building AI agents on Claude and Microsoft Foundry: field notes from e-commerce delivery

A delivery lead's view of model choice, agent tooling and hosting across Anthropic and Microsoft Foundry, including the operating work both approaches require.

Published 2026-07-10 · Updated 2026-09-04 · 7 min · Tom Børvan · https://www.tomborvan.com/writing/building-ai-agents-claude-microsoft-foundry

Most writing about AI agents comes from two places: the developers who build the frameworks, and the vendors who sell them. I sit somewhere else. I run digital commerce delivery at Alpha Solutions Norge, and I put agents into the same projects where I ship storefronts and integrations. The projects I run look like the omnichannel platform we launched for [Møbelringen](https://www.mobelringen.no) in 2024: one of Norway's largest furniture chains, seventy member-owned stores, Adobe Commerce with a Next.js front end on Vercel. Real systems, real users, real consequences when something misbehaves.

From that seat, choosing a model, an agent SDK or framework, and a hosting platform is a delivery decision. Each choice has different implications for identity, permissions, cost and operational ownership. Anthropic's models and tooling can run in several architectures. Microsoft Foundry can provide models, managed agent services and hosting for custom agent applications. The task and the organisation should determine the combination.

## Claude, agent tooling and MCP

Claude's strength, for delivery work, is composability. The agent is a loop: the model reasons, calls a tool, reads the result, decides what to do next. Tools are just functions with a schema. That sounds small. In practice it means I can take a system we already integrate against (a PIM, an order service, a commerce API), expose exactly the operations an agent needs and nothing more, and reason about the blast radius the same way I reason about any integration.

The piece that changed how I scope integrations is the Model Context Protocol. [MCP gives compatible clients a shared protocol](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports) for discovering and invoking a server's tools and resources. One remote server can serve several compatible clients without a bespoke adapter for each one. This site runs a real server at [tomborvan.com/api/mcp](https://www.tomborvan.com/api/mcp).

```text
# MCP endpoint: https://www.tomborvan.com/api/mcp
resources:
  profile://brief   # short bio
  profile://full    # full profile
tools:
  get_profile
  list_projects
  list_timeline
  list_skills
  get_contact_info
```

MCP creates a reusable protocol boundary, while Anthropic's agent tooling can provide the execution loop and tool plumbing. Those are separate choices. Claude Code has also become part of my daily delivery work: reading a codebase, drafting integrations and writing tests. None of these choices removes the need to design permissions, tool behaviour and review paths for the specific workflow.

## Where Foundry fits

[Microsoft Foundry Agent Service](https://learn.microsoft.com/en-us/azure/foundry/agents/overview) offers managed agent types with endpoints, scaling, state persistence and observability. For organisations already operating in Azure, it can fit existing Entra, Azure Monitor and governance practices. The team still has to configure and approve the agent's identity, permissions, data paths, network controls and operating model. Choosing the platform does not supply those decisions automatically.

Microsoft Agent Framework supports custom orchestration and multi-agent workflows, and Foundry can host those applications. Anthropic also documents workflow and agent patterns in [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents). I use several roles only when evaluation shows that the structure improves the result. Model choice and orchestration pattern are separate decisions.

Operations need their own design. [Foundry tracing](https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept) can capture inputs, outputs, tool calls, retries, latency, token use and cost, with traces available in Foundry and Azure Monitor Application Insights. Teams can publish selected business measures in their normal reporting layer, but a dashboard does not define success by itself. The workflow still needs acceptance criteria, review ownership and a response when quality or cost moves outside agreed bounds.

## What project managers get wrong

Here is the mistake, and I have made it. We scope agent projects like software and forget to operate them like operations.

Conventional software lets teams specify many behaviours deterministically. An LLM-backed agent adds probabilistic outputs and variable tool paths. [Agent evaluations](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) therefore need repeated scenarios, task-specific measures and human calibration. Usage, tool and compute costs can grow with each run, alongside any platform licences. "Done" still includes ongoing review after release.

So I still scope the build like software — clear boundaries, defined tools, explicit acceptance criteria — but I plan to verify it like a running operation. Concretely, that changes the questions I ask during delivery:

- How do we sample and review real outputs after launch, not only before it?
- What is the human-in-the-loop path for the cases the agent should not decide alone?
- What does a single run cost, and who watches that number as volume grows?
- When the model changes, how do we notice before the client does?

None of that fits a fixed-scope, sign-off-and-leave contract cleanly, which is precisely why it gets skipped. The teams that succeed budget for the operating discipline up front (evaluation, monitoring, a review loop, a named owner) and treat it as part of the deliverable, not an afterthought the client discovers three months in.

This is why I have stopped framing the decision as Claude versus Foundry. I choose the model, SDK or framework, hosting, identity controls and observability as separate parts of one operating design. The harder question is consistent across stacks: can this organisation evaluate, approve and run the agent after the initial build?

> The model is a component. The discipline to operate it is the product.

[Tom Børvan](https://www.tomborvan.com/#about) · AI Lead & Project Manager, Alpha Solutions Norge. I lead digital commerce delivery and build AI agents for Scandinavian retail.

Exploring an agent pilot for your team? Tell me about the workflow and the people who would use it.

[Discuss an agent pilot](mailto:tborvan@gmail.com?subject=Digital%20commerce%20or%20AI%20project)

---

More writing: https://www.tomborvan.com/writing · Profile: https://www.tomborvan.com/llms.txt · Developer docs: https://www.tomborvan.com/developers