Approach

A practical way to improve AX.

Start with one task. Make the product surface around it clearer. Learn from what happens. Repeat.

Use the brief

Talk about a workflow
01

Observe one workflow end to end

Pick a task a person already wants to delegate, one with a clear goal and a checkable outcome, and watch or replay a single real attempt rather than a demo: what the agent was told, what it could actually see, and what happened when the path was not straightforward.

Record six things while the run happens: the original intent in the requester’s own words, the context the agent actually had rather than the context that should have been available, every tool called and the arguments passed, each point where approval was needed, any retry and what changed between attempts, and the final outcome, including partial completion. A run without this record cannot be compared to the next one.

02

Decompose it with the field map

Once the run is captured, place each moment against the five-system field map on the Agent Experience field map: orientation and discovery, tools and interfaces, workflows and state, delegated action and trust, and evaluation and improvement. Most single workflows touch three or four of the five.

This step turns a narrative into a structural inventory. A moment where the agent chose the wrong tool sits in a different system to a moment where an approval gave no reason, and the two need different fixes even though both can look like the agent simply got stuck.

03

Classify failures against a fixed taxonomy

Not every stall is the same failure. Sort each breakdown against a small, fixed taxonomy rather than inventing a new label for every run: missing or wrong context, an ambiguous or overloaded tool choice, an approval that did not explain the decision, a retry that repeated the same mistake, or a dead end with no route back to a person. Evaluate the decisions in an agent workflow sets out how to judge a decision rather than only the final answer, and Design for recovery describes the recovery patterns a stuck run is usually missing.

A fixed taxonomy is what makes a second review comparable to the first. If every review invents its own vocabulary, nothing accumulates across reviews.

04

Evaluate with a repeatable task set

A single observed run is a starting point, not a result. The next step is a small set of tasks that can be run again, with the model, tool versions, and prompts pinned, so a later run is measuring a real change rather than model or version drift. Build a reliable agent evaluation loop covers how to keep that loop honest, and Measure agent experience sets out what a run record needs to capture.

The task set does not need to be large. It needs to represent the job well enough that a fix which helps this task set is likely to help the workflow generally, and a regression shows up before a person notices it in production.

05

Profile the job with the readiness rubric

Rate the job's six tracks, from finding the page to handing the result back, against the Open Agent-Readiness Rubric and its machine-readable twin at /agent-readiness-rubric.json. Most checks inspect what the product states, such as tool schemas, error payloads or the record left for the person, so the rubric complements the task-based evaluation rather than replacing it.

A stopper or a track at Not yet points at a specific fix. A Dependable profile is not a claim that agents will succeed; Shown means you watched them do it, repeatedly.

06

Produce a recommendation that names the smallest change and its evidence

The review ends with a written recommendation, not a general list of ideas. Name the smallest change that addresses the clearest failure pattern, and attach the run and rubric evidence behind it: which moment broke, which system it sits in, and what the fixed taxonomy called it. The workflow review brief is the template this site uses to keep that recommendation scoped and checkable, and the services page describes how this turns into applied design or engineering work when a team wants help.

A recommendation without evidence is an opinion. Keeping the evidence attached is what lets a reviewer, not only the author, judge whether the proposed fix is the right one.

07

What is deliberately out of scope

The method does not claim client outcomes, benchmark authority, or a ranking against other products. Every source used to shape a guide or a rubric check is public: a specification, a paper, an official document, or a reproducible public method, as described on the about page. Private traces, credentials, and confidential material never enter a public review.

It also does not claim that one review proves an agent will behave the same way in production, on a different model, or on a workflow it has not seen. A review narrows uncertainty about one surface at one point in time. It does not remove uncertainty about everything downstream of it.

What you get

What you get from one review.

A finished review is a short, checkable artefact, not a report of impressions. Applied to a single workflow, it produces:

  • A structural map of one real workflow against the five Agent Experience systems.
  • A short list of failures classified against a fixed taxonomy, not a one-off narrative.
  • A readiness profile for the job, with stoppers listed first.
  • One prioritised recommendation, with the run and rubric evidence that supports it.
  • A repeatable task set you can run again after the change ships.

Limits

Limits of the method.

Treat the method as a way to narrow uncertainty about one workflow, not a way to remove it. It describes declared surfaces and observed failures on the runs actually reviewed; it cannot establish that a fix will hold across every model, every version, or every user population, and it produces no benchmark score that compares one product against another. A task set that passes today can fail again after a model update, a tool change, or a prompt rewrite elsewhere in the system, which is why the evaluation loop in step four is repeated rather than run once. Read a finished review as evidence for one decision rather than a general certificate of quality.

Outcome

A product change that makes the next attempt better.

Good agent-facing systems do not hide complexity. They make the right capability, context, boundary, and route back clear when it matters.

If you want help applying the method, bring the workflow that matters most and the product decision it needs to support.

See the workflow review