Skip to content
All resources
Breakdown For Managing Partner, COO

OpenAI Agent: What It Actually Does – and Whether Your Team Should Care

7 min watch + read Published 24 May 2026 Video companion

Past the demo theatre

What OpenAI Agent actually is

OpenAI Agent is not a chatbot with extra features. It is a system that takes action on your behalf inside a sandboxed virtual environment. Each session spins up a private desktop in the cloud, with access to a browser, a terminal for running code, and connections to external systems such as Google Calendar and HubSpot.

The practical difference from previous tools: the agent does not just produce text in response to a prompt. It navigates, executes, combines outputs, and delivers a finished artefact — a spreadsheet, a presentation, a CRM pipeline report. OpenAI calls this a unified agentic system.

The relevant question for any services leader is not whether that sounds impressive. It is whether the output holds up when your workflows, your data, and your standards are involved. Three practical tests give a useful answer.

Test one

A financial model from public filings

The first test asked the agent to build a discounted cash flow model for a publicly listed company — pulling from regulatory filings, calculating historical and forecast cash flows, running a sensitivity analysis, and delivering a formatted Excel file. The prompt was specific: colour-coded inputs, separate tabs for assumptions, calculations, and outputs, investment-banking-standard formatting.

After roughly eight minutes, the agent delivered a file. It completed the deep research phase autonomously, opened its own terminal to run calculations, and assembled the spreadsheet without manual intervention.

The result was not perfect. Some sections were incomplete and the formatting was partial. But it was a genuine first draft of a complex financial model — one that would have taken a senior analyst several hours to build from scratch. For a Managing Partner or COO who needs a quick-look analysis before a meeting, that is meaningful. It is not a replacement for a qualified analyst producing a board-ready model. It removes the blank-page problem and the hours of data gathering that precede any serious analysis.

Test two

A branded deck from existing data

The second test layered on complexity. The financial data from the first test was uploaded alongside a branded PowerPoint template. The agent was asked to build a ten-slide executive presentation matching the template’s style, incorporating the analysis, and structuring it for a strategy meeting.

Working for thirteen minutes, the agent examined the template, extracted the colour palette and layout conventions, and produced a ten-slide deck that matched the brief. It did not use the actual slide master — it replicated the style visually rather than inheriting it technically. For firms with strict brand governance that distinction matters. For others, a deck that looks right in thirteen minutes rather than three hours is an acceptable trade-off.

The broader point: the agent can chain tasks. Research feeds analysis. Analysis feeds presentation. Each step builds on the last without the human having to pass the baton manually.

Test three

A pipeline review inside a live CRM

The third test moved from public data to a live internal system: log into HubSpot, review the open pipeline, identify overdue deals, and return a prioritised action list. Within four minutes, the agent had reviewed 22 open deals, flagged that every expected close date had been missed by over a year, and produced a structured set of recommendations by deal and by required action.

This introduced a genuine security question, and the answer was two-factor authentication. The agent operates in a browser session, so a compromised password alone is not enough for unauthorised access. The demonstration used a dedicated service account — sensible practice for any AI system operating inside business tools. More broadly, any firm moving agents into live systems needs the Govern step in place first:

  • An approved tool list.
  • Defined data handling rules.
  • A clear policy on which systems AI is and is not permitted to access.

Without that, the question is not whether something goes wrong, but when. This is also the use case with the clearest operational value for a Delivery Director or COO. Pipeline hygiene is one of those tasks everyone agrees is important and almost no one does consistently. An agent that runs that review daily or weekly, without anyone having to remember, addresses a genuine operational drag.

Hype versus gap

Where the tool stops and architecture begins

The vendor noise around AI agents is considerable. OpenAI Agent is not the only tool in this space, and it is not finished. In these tests the Excel model had gaps and the presentation did not inherit the template properly. These are real limitations, not minor quibbles.

The problem with most AI tool announcements is not that the tools are useless. It is that they are presented as complete solutions to problems that actually require architecture, not just access.

Dropping a new AI tool into a team with no documented workflows, no agreed usage policy, and no clear owner for AI outputs does not produce efficiency — it produces noise, inconsistency, and eventually abandonment.

The tool gets blamed. The real issue was never addressed. The demos work because they are well-prompted, well-structured tasks with clearly defined outputs. That structure did not come from the tool. It came from the person who built the prompt.

Read your own position

What this means for a professional services firm

Run the KWA framework against this. The three layers — Knowledge, Workflow, Agents — tell you where you actually are.

  • Knowledge. Does your firm have the data the agent would need, in a format it can access? Financial data scattered across spreadsheets, CRM records six months out of date, and proposals living in personal inboxes are not accessible to an agent. Your information needs to be centralised and structured first.
  • Workflow. Are your processes documented well enough to prompt an agent clearly? Output quality tracked directly to prompt quality. Vague requests produced vague outputs; specific, structured prompts produced usable first drafts. If your team cannot write down what a good output looks like, the agent cannot produce one.
  • Agents. Only once the first two layers are in place does deploying an agent into production make sense. A pipeline review agent reading from a clean CRM with clear deal stages is genuinely valuable. The same agent pointed at inconsistent data returns garbage.

The real starting point

The test worth running

Before evaluating any specific tool — OpenAI Agent or otherwise — the more useful question is: which workflows in your firm are well enough defined that an agent could execute them reliably today?

Most firms, when they answer honestly, find that the number is smaller than expected. Not because their work is too complex for automation, but because the process documentation and data hygiene required to brief an agent properly does not yet exist.

That is the real starting point. Not the tool. The readiness.