Arize Phoenix for Agent Trace Evaluation
Phoenix auto-instruments multi-step AI agents to catch failures that standard monitoring misses.

A single LLM call fails in legible ways: the response is empty, the latency spikes, the API throws an error. An agent pipeline fails in ways that look like success. When a planner, a retriever, a tool layer, and a critic chain together across multiple steps, each one can go wrong on its own terms, and none of those failures resemble a crash. A planner can hallucinate a step that was never grounded in the task. A retriever can return context that's stale by months but syntactically indistinguishable from fresh data. A tool call can fail, and it can get swallowed without ever raising an exception the pipeline recognizes. A critic meant to catch mistakes can overcorrect and discard a correct answer. Each of these produces different evidence: wrong text, missing data, a misleading success code, an action taken without the authority to take it. A standard error monitor, built to watch for exceptions and status codes, catches none of it.
Microsoft's updated taxonomy of agentic AI failure modes names this problem directly, introducing categories such as goal hijacking and agentic supply chain compromise. These failures are not detectable in a single LLM call examined in isolation. They emerge only when an agent operates autonomously across a sequence of steps, where each step trusts the output of the one before it. That chain of trust is what makes the failures hard to see: an agent can return an HTTP 200 while fabricating identifiers that look legitimate, executing an irreversible action outside its policy boundaries, or reporting a status message that misrepresents what it actually did. A database can be altered, a record can be invented, and the system logs will show nothing but a clean, successful run. This is the gap that agent observability has to close, and it's a different problem from watching whether an LLM API responded correctly. It requires seeing the full shape of a multi-step run: what each component did, what it passed to the next one, and whether the final claim of success matches the actual state of the world.
Phoenix's design and structure
Phoenix is an open-source observability and evaluation platform, and it's built for LLM and agent applications. It's built on OpenTelemetry and the OpenInference semantic convention, and it can be self-hosted without a vendor API key. As listed in the Awesome AI Agents index, Phoenix covers tracing, evaluations, datasets, experiments, prompt management, and debugging workflows, all built on those same open standards.
Phoenix sits at one end of a two-tier structure. Inside the broader Arize ecosystem, Phoenix is the open-source entry point for building and testing AI systems, while Arize AX extends the same underlying standards and workflows into a managed platform meant for production-scale observability, online evaluations, monitoring, and continual improvement. Phoenix is licensed under the Elastic License 2.0, which allows a team to run it on its own infrastructure but restricts offering Phoenix itself as a hosted, managed service to others. The detail that matters most for everything described later is the OpenInference spec. Because every span Phoenix records carries a typed role, whether it's an LLM call, a retrieval step, or a tool invocation, Phoenix's evaluators always know what kind of thing they're looking at. That typing is what lets Phoenix score a trace without a team first having to hand-label what each piece of it represents.
How Phoenix captures an agent run as a trace
After instrumenting an agent, the Phoenix UI opens to a list of projects and traces, each with a latency and cost summary attached. Click into one and it opens into an execution tree showing the full path the run took, from agent to LLM call to tool invocation. Drill further and the prompts, the tool arguments, the raw outputs, and the final answer are all there to read.
Getting to that view doesn't require writing manual span code for every operation an agent performs. Phoenix's OpenInference wrappers auto-instrument popular frameworks, so a team can wire a running agent into Phoenix with one registration call. The register() helper, part of phoenix.otel, points a tracer provider at the Phoenix collector, and framework-specific instrumentors then patch the framework so that every call it makes emits a span on its own. Each span carries its inputs, its outputs, its latency, its token usage, any errors it hit, and its typed role under the OpenInference schema. Related spans roll up into a single trace with clear parent-child relationships, so a multi-step run reads as one connected structure, not a pile of disconnected logs.
Oracle's integration of its Open Agent Specification with Phoenix shows what that consistency buys in practice. Because the trace format stays the same across different runtimes, such as LangGraph and WayFlow, dashboards and evaluation setups don't have to be rebuilt every time a team swaps out a runtime or an underlying model. When systems are built from multiple agents working together, Phoenix goes a step further and abstracts the raw spans into a node-based graph that maps how agents and tools call on each other. That graph view matters most when the question isn't why one call failed but why an orchestration of several agents produced the wrong outcome. Teams running custom steps, such as a bespoke retrieval method or internal business logic that no auto-instrumentor covers, can decorate those spans manually and still get the same nested trace tree as everything else.
Scoring traces with Phoenix's built-in evaluators
A trace on its own tells a team what happened. It does not tell them whether what happened was good. Phoenix closes that second gap with its evaluation functions. The run_evals() function, or the lower-level llm_classify for more control, runs against a DataFrame of spans and writes its results back onto them: a label and an explanation attached to each one. That write-back is what makes the loop useful. A team doesn't run an evaluation off to the side and then try to reconcile it with the trace later. Every span simply gets a verdict, such as a hallucination flag of yes or no, sitting right next to the prompt and completion that produced it.
Phoenix ships a catalog of general-purpose evaluators, including HallucinationEvaluator, QAEvaluator, and RelevanceEvaluator, each built on a tested rubric with an LLM acting as judge. MCP tracing support, built on the same open OpenInference specification, extends this same coverage to tool calls made through MCP-connected coding agents, so those calls get scored the same way any other span does. For criteria no general evaluator covers, such as a compliance rule, a citation style, or a tone requirement specific to one domain, llm_classify takes a custom prompt template, a defined list of allowed output labels (the "rails"), and the spans DataFrame, and returns a labeled DataFrame built to that rubric.
What this produces in daily use is a workflow some teams call filtering to failure. Once evals have run, a team sorts and filters traces by outcome, jumps straight to the spans that failed, reads the prompt and completion that produced the failure, and reads the judge's explanation for why it was marked wrong. All of it happens in one view, without leaving Phoenix to cross-reference a separate log system or spreadsheet.
Using datasets and experiments to catch regressions before shipping
Scoring a trace after the fact tells a team about one run. Catching a regression before it ships requires comparing many runs against a fixed standard, and this is what Phoenix's datasets and experiments model is built for. A dataset in Phoenix is a fixed collection of input and output examples that acts as a stable test surface. Production traces can be promoted directly into a dataset, so the regression suite a team builds reflects real failure cases pulled from actual runs, not synthetic examples invented in a vacuum.
Against that fixed set, run_experiment() runs the application's task function with evaluators attached, then it stores both the run and its evaluation results in the Phoenix database. Running it again after changing a prompt or swapping a model produces a direct score comparison over identical inputs: a prompt change or model swap either holds quality steady or it doesn't, and the comparison shows which. The same harness carries across runtimes, too. As the Oracle integration with Agent Spec demonstrates, running one agent on multiple runtimes and comparing the outcomes uses a single evaluation harness without touching the instrumentation underneath it.
The most concrete version of this appears in continuous integration. A team can run Phoenix as a container inside a test job, execute run_experiment() over a regression dataset as part of that job, and assert on the aggregate score. If a pull request drops QA correctness below a set threshold, it fails the build, the same way a broken unit test would. Curating and maintaining a dataset takes effort, and teams under shipping pressure often don't make time for it, but Phoenix lowers that cost by letting traces from production feed the dataset directly. The test corpus grows out of what the agent actually encountered, not out of a separate authoring task someone has to schedule.
Phoenix in the development workflow: the MCP server and coding agent integration
Phoenix's MCP server connects coding agents, such as Cursor or Claude Code, directly to a Phoenix deployment. Through that connection, an engineer gets tracing, datasets, experiments, evals, and prompt workflows inside the same interface already used to write and edit code. Trace inspection, dataset access, experiment runs, eval results, and the prompt playground are all reachable through MCP without switching over to the Phoenix UI for each step of debugging.
Phoenix is actively extending test coverage for its own MCP-exposed skills, using Claude plugin evals to check skill triggering and Harbor for stateful, multi-step workflows that need a real Phoenix server, persisted data, tools, or environment verification to run correctly. The direction behind it is straightforward: observability moves into the environment where engineers are already working, rather than requiring a separate monitoring console they have to remember to open. As coding agents increasingly become the primary way engineers write and debug agent code, tools like Phoenix need to live there to stay useful.
Where Phoenix's coverage ends
Phoenix captures what an agent did with real precision, and it scores whether the output was good against the rubrics a team defines. Several reliability questions sit outside that scope by design, and some require Arize AX, the paid tier beyond self-hosted Phoenix.
Context correctness is one persistent gap. Tracing shows what happened, and evals show whether the output met a defined bar, but neither one confirms whether the material an agent retrieved was current, complete, and authorized to use. That matters most for agents pulling from live data sources, where the answer can read as fluent and well-formed while resting on information that's already out of date. Continuous production monitoring is another boundary. Online evals and drift or bias monitoring against live traffic sit in the AX tier, so they aren't part of self-hosted Phoenix, though Phoenix does include batch-mode drift detection and embedding drift visualization. A team running open-source Phoenix gets batch evals and experiment-based regression testing, not a persistent watch over production traffic as it happens.
Domain-specific judgment calls are left to the team as well. Whether a medical triage was clinically sound, or whether a legal brief held up jurisdictionally, requires a custom eval written through llm_classify. Phoenix supplies the mechanism for running that evaluation at scale, but it doesn't supply the rubric; someone on the team still has to define what "correct" means in that domain. The same is true of failure-mode coverage more broadly. Phoenix's four built-in agent evaluators, covering function calling, path convergence, planning, and reflection, handle structural agent behavior well, but a team still has to decide which failure modes matter for its own use case and build or configure evals around them. Phoenix doesn't go looking for failure patterns no one asked it to check. And slow, silent behavioral drift, an agent gradually shifting its behavior in ways no single trace flags as wrong, needs aggregate pattern analysis across many traces over time, which sits above what span-level evals catch one at a time.
None of this is a flaw in Phoenix so much as a boundary of what a tracing-and-evaluation platform is built to do. Knowing where that boundary sits is what lets a team decide what else it needs to put in place around it.
Fitting Phoenix into a production reliability strategy
Phoenix belongs in most agent reliability stacks as the instrumentation and evaluation layer: the place where a team sees what an agent did, step by step, and scores whether each step held up. It fits naturally in the inner development loop, running locally, in a notebook, or inside CI; in regression gating against a dataset; in structured debugging of an individual trace; and anywhere OpenTelemetry compatibility matters because the stack spans multiple frameworks or runtimes.
The failure modes that do the most damage in production, though, are the ones that need something more than that. Silent behavioral drift, a policy violation buried in an otherwise clean-looking run, a fabricated output riding on top of a success code: these require a layer that understands what an agent was supposed to do, not only a record of what it did. Phoenix's evaluators are only as good as what a team tells them to check for. A failure mode nobody thought to write an evaluator for will pass through unscored, no matter how detailed the trace underneath it is. That's the real argument for treating Phoenix as one layer in a broader reliability strategy rather than the whole of it: it gives a team a precise, repeatable way to see and score what it already knows to look for, and the rest of the strategy has to cover what it doesn't yet know to ask.