Arthur AI provides tools to discover, evaluate, observe and govern AI applications. Its current platform emphasizes agents: which ones exist, how they use models and tools, and whether their behavior meets an organization’s policies. The central implementation detail is easy to miss in a product demonstration: Arthur can return a guardrail verdict, but the application must be wired to enforce it. A logged violation is not the same thing as a blocked action.
- 01Best fit Teams that need model evaluation and agent-governance evidence connected to operational workflows.
- 02Architecture choice Open-source evaluation components can run beside workloads; managed platform and enterprise options add distinct capabilities.
- 03Evaluation boundary Sequenced reviewed public documentation and pricing, without operating an Arthur deployment.
01 / ProductFrom application traces to policies with accountable owners
The current platform combines agent visibility, evaluation, runtime-security claims and cost analysis. Arthur’s scope is broader than checking a final answer for hallucinations. A tool-using system can retrieve the wrong document, call an inappropriate tool or spend excessively while still producing a plausible final response. Traces provide a place to investigate those intermediate steps.
The agent-governance page describes several discovery approaches, including OpenTelemetry streams, MCP-server monitoring, network analysis, cloud APIs and endpoints. These address different ways agents enter an organization. The customer still needs to establish which sensors are installed, what permissions they receive and what part of the environment remains outside their view.
The Arthur Engine repository covers guardrails, evaluation, tracing, prompt management and experiments. It distinguishes functionality available through the engine from capabilities requiring the Arthur Platform. That matters for self-hosting decisions: an open-source component is not automatically a self-hosted copy of every feature shown in an enterprise sales presentation.
02 / AudienceFor builders and governance teams who share the same evidence
A product team may want to understand why a support agent’s answers deteriorated after a retrieval change. A security team may want to know whether the same agent exposes customer data. A governance lead may need to show who approved its use. Arthur is relevant where these questions should refer to the same application identity, version and observed behavior.
A team with one low-impact drafting tool may not need organization-wide discovery. It can start with a narrow evaluation and tracing route if those solve the immediate problem. At the other extreme, a large organization should avoid treating a dashboard as governance by itself: someone must own policies, review exceptions and act on findings that cross departmental boundaries.
Our Arize AI blueprint explores observability and evaluation around AI applications. Our Fiddler AI blueprint examines model monitoring and explainability. These comparisons are useful when deciding whether the principal problem is investigating quality, understanding model behavior or connecting many agents to a broader governance process.
03 / WorkflowProposed workflow for a support assistant with sensitive context
This is a proposed evaluation, not hands-on testing by Sequenced. Choose a support assistant operating on synthetic customer records. Define the permitted job: retrieve the requesting customer’s order status and draft a response. It should not reveal another customer’s information or execute an account change. Establish these rules in the application before evaluating how Arthur observes and checks them.
Deploy the relevant engine and connect the application using the official setup route. The site describes Docker-based operation and points to deployment resources. Keep the evaluation environment separate from production, record the engine version and ensure that any model-provider credentials belong to the test environment. The pilot should not need access to real customer records.
Instrument the retrieval and tool steps so a trace can show which customer identifier was used and what information was returned. Record the model and prompt version alongside the request. When an answer is wrong, the team should be able to distinguish a retrieval mismatch from a model interpretation error rather than attributing every failure to the language model.
The platform quickstart demonstrates setting up a prompt-injection metric and inspecting results in a playground. That page is older than the current marketing surface, so use it as an implementation example and verify the current onboarding flow. It supports the mechanics of a pilot, not a claim that one demonstrated metric covers every type of attack.
Configure the application to act on the returned verdict before sending a response or invoking a sensitive tool. Arthur’s governance explanation explicitly assigns blocking, redaction or re-prompting to the application’s integration. Test a forbidden request and confirm that the downstream action did not happen. This is the difference between measuring a violation and preventing its effect.
Run a legitimate request containing security-related language as a false-positive check. Then change the retrieved document to include a harmless conflicting instruction and observe whether the result is flagged and enforced. Keep the content, rule version and actual application outcome together. A correct detection with a broken enforcement branch is still a failed control.
Finally, introduce a controlled instrumentation gap in the test environment. Confirm whether the platform identifies that a policy is no longer being applied and how the owner is notified. Assign the finding to a person who can repair the integration. This tests the durability of governance after a deployment change, rather than only the initial successful setup.
04 / PricingPublic tiers exist, but enterprise scope needs confirmation
The pricing page displays Free at $0 per month, Premium at $60 per month and Enterprise with custom pricing. It presents Free monitoring for up to four use cases and Premium for up to one hundred, alongside other limits. The page uses dollar signs without an explicit currency code in the reviewed copy; confirm currency and any applicable taxes when budgeting.
The detailed comparison also lists project, retention and usage allowances. Free and Premium are not simply unlimited versions of the same operational scope. In particular, the page lists seven-day retention for Free and thirty-day retention for Premium. A pilot may fit those windows while an organization’s incident-review needs do not.
Arthur separately offers an open-source Evals Engine. Its repository carries an MIT licence, but operation still requires infrastructure and, for relevant evaluations, model access. Treat the engine’s software licence, platform subscription and underlying inference costs separately. Enterprise discovery, deployment and security requirements should be confirmed against a proposal rather than inferred from the lowest public tier.
| Route | Displayed basis | Relevant boundary |
|---|---|---|
| Free | $0 per month | Up to 4 monitored use cases; 7-day retention. |
| Premium | $60 per month | Up to 100 monitored use cases; 30-day retention. |
| Enterprise | Custom | Confirm deployment, SSO, support and custom usage. |
| Open-source engine | MIT-licensed repository | Hosting and model calls remain operational costs. |
Public tiers from Arthur pricing, consulted 11 October 2026. Dollar amounts are displayed without an explicit currency code; enterprise entitlements and third-party model costs require separate confirmation.
05 / DistinctionsApplication-side enforcement makes the control inspectable
Arthur’s explicit separation between evaluating and enforcing a request is useful architecture information. It allows a team to inspect exactly where a verdict becomes a block, redaction or escalation. It also creates a responsibility: if the caller ignores the verdict or uses a different path, a correct evaluation may have no protective effect.
The platform’s proposed hybrid architecture keeps an evaluation data plane near workloads while a central control plane receives selected information. That can help organizations reason about data locality, but the data actually exported must be verified in their integration. “Local evaluation” does not automatically settle what appears in traces, support bundles, external judge-model calls or aggregated reports.
Bringing traces and evaluations together can make a regression easier to diagnose. For the support assistant, a low groundedness result should lead to the exact retrieval and response that produced it. The operator can then ask whether the source was missing, the query was wrong or the model ignored the result. This is more actionable than a weekly quality average without examples.
06 / QuestionsAsk what is observed, retained and actually stopped
Discovery coverage depends on the configured sensors and the environment they can reach. Verify at least one known agent on each required surface and keep a record of exclusions. An agent found through a network signature may have less detailed evidence than an instrumented application. The platform inventory should preserve that difference rather than implying identical visibility.
For guardrails, test timeouts and unavailable dependencies. Decide what the application does if a verdict never arrives, and verify that retries cannot repeat a sensitive operation. Also measure legitimate requests incorrectly blocked. A control that operators routinely bypass because it disrupts normal work is unlikely to remain reliable.
For commercial evaluation, map the pilot’s traces, evaluations, inferences and retention needs to the actual plan definitions. Several counters can grow from one user request. Obtain the relevant overage and retention terms before extending the integration to high-volume traffic, and verify that any required enterprise deployment or support arrangement is part of the order.
07 / DecisionProve one policy from request to observed outcome
Arthur is a useful candidate for organizations wanting evaluation and governance grounded in application evidence. The first meaningful result is a request whose path can be explained: which agent handled it, which data it used, which policy evaluated it and what the application actually did. Expand only when that chain works for both permitted and prohibited behavior.
Need to debug one unreliable assistant
Instrument its retrieval and tool steps, attach evaluations and inspect complete failure examples.
Need enterprise agent governance
Test discovery coverage and one enforceable policy across the actual application path before expanding the inventory.
Require local data processing
Evaluate the engine and proposed deployment while tracing every export and external model call.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- Arthur platformConsulted
- Agent security and governanceConsulted
- PricingConsulted
- Evals Engine setupConsulted
- Arthur Engine repositoryConsulted
- Platform quickstartConsulted

