Microsoft’s Copilot bet: a toolbox of models, not one “best” model | The Neuron

Microsoft’s Copilot bet: a toolbox of models, not one “best” model

GPT-5.2 is now available in Microsoft 365 Copilot and Copilot Studio; part of Microsoft’s push to deliver new frontier models on launch day while expanding customer choice.

Written By
Corey Noles
Corey Noles
Dec 15, 2025
6 minute read

When Charles Lamanna talks about Microsoft’s Copilot roadmap, he doesn’t start with a single model, a single benchmark, or a single “winner.”

He starts with a toolbox. That’s a subtle but important tell.

Lamanna is Microsoft's President of Business Apps & Agents and he's tasked with ushering in the new era of workplace AI.

AI isn’t fighting a cage match where one model emerges and everyone standardizes forever. It’s a world where models are getting better in uneven, “jagged” ways—meaning the best model for this job might be the wrong model for that job. And the product that wins isn’t the one with the flashiest demo. It’s the one that can match the right capability to the right task, inside the tools people already use, grounded in the data they already have.

In other words: the Copilot story isn’t “Model X is here.” The bigger story is Copilot becoming a multi-model work layer—routing requests, orchestrating tools, grounding answers in your work context, and (importantly) letting a small set of power users override defaults when it matters.

He also noted that when it comes to Copilot, speed of access matters. He noted GPT-5.2 was available in their tool on launch day, so it's available to users when they want it. In the AI space, they generally want it yesterday. The company plans to strive for that approach moving forward, as well.

Why Microsoft is leaning into “model choice” now

Not long ago, the standard product instinct was: users don’t want to pick models. They want the system to “just work.” That was a broad assumption no more than 12 months ago widely held by non-technicals and business leaders.

Lamanna says that assumption has cracked for two reasons.

First: users got smarter. A lot smarter. People now notice that prompts behave differently across models. They can feel strengths and weaknesses. They also have opinions. In Lamanna’s framing, “user choice has become an important thing” because users now expect it and it’s useful.

Advertisement

Second: the models stopped moving in a straight line. We’re not in a world where “the newest model is best at everything.” Some models pull ahead in coding. Others do better at broad reasoning. Others handle work-context tasks more reliably. The “one best model” idea held up better a year or two ago. Today, it’s increasingly a myth.

So Microsoft’s stance becomes a balancing act:

  • Default to “auto” routing for most people (because nobody wants a settings menu to do their job).
  • Offer explicit model choice for the influencers, the “2%” power users, because they shape adoption across teams.

And here’s the key detail: those power users don’t want to leave work to get that control. They want to choose the right model without losing the Microsoft 365 context—their files, meetings, emails, Teams threads, and all the organizational memory embedded in them.

This is what “model choice” looks like when it’s actually a product strategy: control where it matters, invisibility where it doesn’t, and context everywhere.

The real unlock: Tool use and long chains of work

If model choice is the strategy, tool orchestration is the payoff.

Lamanna describes Microsoft 365 Copilot as “extremely heavy in tool use,” because the whole point is to do real work: pull the right file, reference the right thread, check the calendar, summarize the meeting, draft the follow-up, update the doc, and keep moving.

Also, models are getting better at sequencing multi-step tool calls without breaking the chain.

A big difference from 9–18 months ago is reliability across steps. Lots of early agent demos died in the middle of the workflow: wrong parameter, wrong tool, wrong file, wrong turn. Better models (and better scaffolding) reduce that failure rate. And that’s where you start seeing “emergent” scenarios, like use cases that don’t just look good on stage, but hold up in messy real life.

Lamanna’s most concrete example is the kind of detail that should be in more AI coverage: in Researcher-style experiences, newer models can string together 30–40 tool calls to produce a single response (and sometimes more).

You don’t need to care about model lineage or naming conventions to care about this:

  • Single-step AI saves time on drafts.
  • Multi-step AI changes workflows.
Advertisement

“Swap the model, keep the work brain”

Multi-model sounds chaotic until you realize Microsoft is trying to standardize the layer around the model.

Lamanna points to Microsoft’s work-context layer, a tool called WorkIQ, as the glue that makes the approach coherent. The idea is straightforward: Microsoft can “swap the model” but keep the context, retrieval, permissions, tooling, and enterprise scaffolding.

This is Microsoft’s real bet:

  • Models will keep changing fast.
  • The durable advantage is the system that connects models to work.

In that framing, GPT-5.2’s arrival matters (and it does), but mostly because it plugs into the same work layer where the real product differentiation lives.

Trust and retrieval: “Readjust your priors”

Lamanna’s advice on trust is blunt: if your opinion of AI was formed 18 months ago, use it again and “readjust your priors.”

He points to three changes that (in combination) move Copilot-style systems closer to something you can rely on:

  1. Models are improving on hallucinations and tool execution.
    Not just “smarter,” but better at the unglamorous part: choosing the right tool and filling the parameters correctly.
  2. The experience is increasingly built around receipts.
    Citations, sources, and grounding are newer mechanisms that let a user verify where the answer came from, and click through when it matters.
  3. Behind-the-scenes retrieval is getting heavier.
    Better latency and performance means you don’t have to rely on one query. You can do many. Lamanna describes scenarios where the system runs something like 20 search queries that can be cross-checked and synthesized, reducing the odds that one bad retrieval poisons the output.
Advertisement

Adoption is a product problem, not an awareness problem

One of the most practical insights here has nothing to do with models.

Lamanna says Microsoft learned a simple lesson: if you have great AI but nobody uses it, you get no value. And changing habits is hard because people are busy.

So Microsoft’s adoption strategy is interface-first:

  • Put Copilot where people already work (Word, Excel, Teams).
  • Reduce the friction of “going somewhere else to do AI.”
  • Use those in-app moments to teach users what’s possible.

At the same time, Lamanna describes a balancing act: keep “AI-ifying” every core app so it doesn’t feel dated, while still pulling users toward the Copilot app as a more unified “AI-first” starting point.

This is the unsexy truth of enterprise AI: the model can be brilliant, and still lose to the calendar.

How customers should measure Copilot agents

If you’re writing for people who actually have to justify this stuff (read: most of the planet), Lamanna’s measurement ladder is the right way to close:

Level 1: Usage
If nobody uses Copilot or an agent, there’s no value. Period.

Level 2: Outcomes tied to business processes
Pick workflows that map to measurable outcomes (like HR deflection (tickets solved without escalation), sales conversion, finance reconciliation, procurement automation) then pair them with a quality metric so you’re not “optimizing for nonsense.”

Level 3: P&L impact
The boldest level: does this show up in dollars and cents? Not “we feel faster,” but “revenue per seller is up,” or “the growth curve changed.”

He also flagged a trap that every exec should tattoo on their forehead: quantity metrics alone can be gamed. You can get to “100% automation” by having the agent say no to everything and block escalation. That looks great on a dashboard and terrible in real life. The fix is simple: pair volume with quality like satisfaction scores, resolution correctness, downstream impact.

Advertisement

Where this leaves Copilot

Microsoft’s message, through Lamanna, is basically this:

  • Stop thinking of Copilot as “a model inside Office.”
  • Start thinking of it as a work system with models + routing + tools + context + UI + measurement.

GPT-5.2 joining Copilot is part of that story, but it’s not the whole story. The bigger shift is Copilot turning into an agent platform that can actually run multi-step work reliably, while giving power users just enough control to pull the organization forward.

The workplace AI race is starting to look less like “who has the smartest brain” and more like “who has the best nervous system.”

And Microsoft is betting that the nervous system wins.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.