ZenML helps teams keep the inputs, execution and outputs of AI work connected. Its pipeline framework turns Python steps into reproducible runs on a chosen infrastructure stack; Kitaru adds a separate way to investigate and replay agent sessions. The company is relevant when a successful notebook or agent demonstration has become a changing production system, and the team needs to explain what changed before shipping the next version.
- 01Two related products. ZenML orchestrates pipelines; Kitaru evaluates agents through recorded sessions and replay. Neither requires adopting the other.
- 02Infrastructure stays explicit. Stack components determine where orchestration and artifacts live; managed metadata does not remove responsibility for the underlying systems.
- 03Different billing units. ZenML managed tiers use pipeline executions, while Kitaru has a separate workspace price and capacity limits.
01 / ProductPipelines and agent replay under one company
The company overview presents ZenML and Kitaru as two products available under the Apache 2.0 open-source licence. ZenML is the orchestration foundation for machine-learning and LLM pipelines. Kitaru addresses a different problem: turning production agent sessions into evidence for improving behavior. Keeping those responsibilities distinct prevents a buyer from confusing a successful pipeline run with a successful agent decision.
The ZenML product page describes versioned step outputs, caching and interchangeable infrastructure components. A trained model, an evaluation dataset and a generated report can become related artifacts rather than disconnected files. That relationship is valuable when the next person investigating a regression did not write the original training code.
The core concepts guide organizes work around steps, pipelines, artifacts and stacks. These are operational records and execution abstractions. They do not automatically establish that an input dataset is representative or that a score reflects the business outcome the team actually cares about.
02 / AudienceTeams whose experiment history must survive a handover
A good candidate is an ML team with existing cloud storage and compute that wants consistent lineage across projects. It may already know how to train a model, but struggle to reconstruct which data and dependencies produced the deployed version. ZenML is useful when replacing that informal knowledge with inspectable run records matters more than choosing another model algorithm.
A second candidate operates an agent and already collects traces, yet changes prompts through manual trial and error. Kitaru provides a concrete evaluation route for that situation. A team without production sessions can still begin with examples, but should not mistake a tidy example collection for the distribution of tasks users will eventually send.
The Databricks blueprint is a useful comparison when the desired purchase includes a broader data and AI platform. The Arize blueprint provides context for observability and evaluation. ZenML's decision is more specifically about repeatable execution and, through Kitaru, testing a change against recorded agent behavior.
03 / WorkflowA proposed release loop for a classifier and its support agent
Consider a product that classifies incoming documents and uses an agent to explain uncertain cases to an operator. The following workflow is a proposed design based on public documentation, not a system tested by Sequenced. Start by defining the document population, the held-out evaluation set and the exact conditions that should send a case to human review.
Build separate pipeline steps for data preparation, training, evaluation and report generation. Give the inputs stable versions and save outputs as artifacts. A report should identify the candidate model and the evaluation data it used. Otherwise, a seemingly improved score may simply reflect a different test set, and the run history will preserve an ambiguous comparison.
Choose the execution environment using the stacks guide. An orchestrator and artifact store are foundational components; other integrations extend the stack. A local development run and a cloud training run can share pipeline logic while using different components. Verify access to each artifact store from the selected execution environment before assuming a local success will transfer.
Enable caching only where the inputs accurately describe the work. Reusing deterministic preprocessing is attractive; reusing a result from a mutable external service can conceal a change. Record a dataset version or another explicit freshness input when upstream state matters. The practical test is whether an engineer can explain why a step reused its output instead of executing again.
For the agent, use the Kitaru quickstart to distinguish importing existing traces from instrumenting code for replay. Imports make historical sessions available for investigation. The adapter supplies the connection needed to rerun the agent's code against recorded tool responses. Having a trace file alone is not the same as having a complete executable replay environment.
Select a cohort that includes successful explanations, unsupported claims and cases correctly escalated to an operator. Turn an accepted failure into an evaluator with an explicit pass condition. Then change one prompt or tool policy and compare the candidate with the baseline. Include difficult cases that the change was not designed around, so a narrow repair does not hide a new regression.
The Kitaru product page emphasizes replaying against recorded context. Treat that as a controlled experiment. A recorded service response can help isolate a code change, but it cannot prove how a live service will behave tomorrow. Follow the replay with a small, observed deployment before expanding the change to every production session.
04 / PricingManaged pipeline executions and agent replay have separate prices
The pricing page, consulted 7 October 2026, lists free self-hosted options and paid managed workspaces. The displayed ZenML Scale selection is $999 per month for 2,000 executions, three projects and five snapshots. Its slider changes the execution tier, so that figure should be treated as the captured selection rather than a universal starting price.
Kitaru Cloud is separately displayed at $39 per month with three agents, two seats and 90-day session retention. The page describes one subscription and bill while distinguishing the two pricing models. Confirm the combined order, workspace scope and required governance features before budgeting both products. These displayed dollar fees should not be read as covering the cloud compute or model-provider usage in the proposed workflow.
| Route | Commercial basis | Decision boundary |
|---|---|---|
| ZenML Open Source | Free framework; self-hosted | Operate infrastructure and integrations yourself |
| ZenML Scale, displayed selection | $999/month; 2,000 executions | 3 projects and 5 snapshots; tier slider changes scope |
| Kitaru Cloud | $39/month | 3 agents, 2 seats, 90-day retention |
| Enterprise | Custom quote | Confirm governance, hosting and workspace scope |
Selected ZenML and Kitaru pricing, consulted 7 October 2026. Displayed dollar amounts; infrastructure and model usage require separate budgeting.
05 / DistinctionsThe useful connection is between execution history and decisions
ZenML's infrastructure abstraction can reduce the amount of pipeline code tied to a particular orchestrator. The benefit is practical portability of a workflow definition, not evidence that moving a production system requires no engineering. Container images, credentials, networking and artifact access still need to work in the destination. A realistic migration test includes those surrounding dependencies.
Kitaru adds a different kind of continuity: preserving a useful failure as an evaluation case instead of letting it disappear into an incident discussion. That can make an agent improvement review more concrete. The team still supplies the judgment about what constitutes a good answer, and should examine disagreements between automated evaluation and domain reviewers rather than averaging them away.
Together, the products support a useful distinction between repeatability and quality. A pipeline can consistently reproduce a weak model. An evaluator can consistently reward the wrong response. The stronger operating process links reproducible artifacts to a clearly defined acceptance decision and retains enough evidence for another engineer to challenge that decision later.
06 / QuestionsResolve what reaches the control plane and what replay can reproduce
For a chosen deployment, map the locations of metadata, artifacts, recorded sessions and replay workers separately. ZenML's metadata-layer description should not be stretched into a blanket statement that every Kitaru trace stays outside a hosted service. Session records can contain sensitive prompts and tool results even when the worker and model credentials remain in the customer's environment.
Check replay coverage for the actual agent framework and tool patterns. An imported session with missing tool output may be useful for review but insufficient for the intended experiment. Ask how unsupported calls, changing external state and repeated side effects are handled. Use a test account and controlled tools for the first replay rather than connecting a write-capable production workflow without examining its behavior.
This review used public product and documentation sources. We did not execute ZenML pipelines, inspect a private Pro workspace or measure Kitaru's ability to reproduce a particular production agent. The next useful evidence is a complete run that another team member can reconstruct, plus an evaluation change whose acceptance criteria were agreed before seeing the result.
07 / DecisionStart with a change that somebody else must be able to explain
Choose ZenML when the immediate problem is preserving reliable execution history across an evolving AI workflow. Choose a Kitaru pilot when agent traces exist but the team lacks a repeatable improvement loop. The best first milestone is modest: reproduce one pipeline result or demonstrate one agent change against a declared cohort, with the data, code, cost and approval decision all identifiable.
Reproduce one release candidate
Connect a fixed dataset, chosen stack and model artifact to a report another engineer can rebuild.
Replay one consequential change
Use a declared cohort and evaluator, then inspect improvements and regressions before a limited rollout.
Validate the integration boundary
Test artifact access, metadata exposure and the actual managed tier before moving several teams.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- ZenML company overviewConsulted
- ZenML productConsulted
- Kitaru productConsulted
- ZenML core conceptsConsulted
- ZenML stacksConsulted
- Kitaru quickstartConsulted
- ZenML pricingConsulted

