sequenced.ai
Articles/Workflow & automation/Blueprint//8 min read

Union.ai connects durable AI tasks, model artifacts and serving in your cloud

Understand Union.ai and Flyte, deployment responsibilities, usage meters and a proposed model evaluation workflow with traceable artifacts.

By Sequenced deskAI-assisted, source-led · how we work
Visit Union.ai website ↗
FlyteOpen-source foundationPython tasks and durable execution
TaskEnvironmentExecution definitionImage, resources and secrets
ArtifactsVersioned outputsNamed references with lineage
AppsServing layerLong-running APIs and model endpoints
Union.ai mark
Union.aiunion.ai · independent research

Represent this company? Verify your work email to access its workspace, or send the desk a factual correction.

Union.ai provides a runtime for AI workflows whose work spans multiple tasks, machines and failure conditions. Built around the open-source Flyte project, it connects Python execution with versioned artifacts and services running in a customer's environment. It is most relevant when training, evaluation or agent workloads have grown beyond a single job and need a durable, inspectable path from inputs to an approved output.

In brief
  1. 01A company around Flyte. Union.ai supplies the commercial platform; Flyte is its open-source execution foundation.
  2. 02Infrastructure ownership varies. Self-managed and managed BYOC deployments both keep the data plane in the customer environment, with different operating responsibilities.
  3. 03Usage has several meters. Actions and allocated compute contribute to platform usage; optional managed deployment has a separate fee.

01 / ProductA runtime joining batch work to the services that use its outputs

The Union.ai user guide describes a platform for model factories, production agents and model serving. The common substrate is durable Python work that can be recorded and recovered. The platform is not a model provider: the application still chooses the model, supplies its data and defines what an acceptable result looks like.

The task guide defines a task as a Python function running remotely in a container. A TaskEnvironment declares shared execution requirements such as the image, resources and secrets. Tasks can call other tasks, allowing the graph of work to follow ordinary program control flow rather than requiring every branch to be drawn separately in advance.

The app guide covers long-running services such as dashboards, APIs and model endpoints. This provides a way for a task's output to become something another system can consume. A service remaining reachable is a different operational requirement from a training task completing, so the team should examine both execution histories and serving behavior.

02 / AudienceTeams with expensive work that must be recoverable and explainable

Union.ai is relevant to an AI platform team coordinating repeated training and evaluation jobs across an existing cloud environment. It can also suit an agent product where a single user request fans out into several long-running operations. The common requirement is understanding what already completed, what needs to retry and which output belongs to the current version of the workflow.

The platform demands clear task boundaries and infrastructure ownership. A group with one occasional script may find that a full distributed runtime adds more work than it removes. A stronger candidate already has meaningful failure modes: interrupted GPU jobs, duplicated external work, partial batch completion or difficulty tracing a deployed model back to its evaluation record.

The Temporal blueprint is a useful comparison for durable application execution. The Anyscale blueprint provides context for distributed AI compute around Ray. Union.ai's decision centers on how Flyte tasks, artifacts and services fit a production AI lifecycle. These systems can address adjacent layers, so compare a complete workload rather than assuming their feature lists describe identical responsibilities.

03 / WorkflowA proposed evaluation gate before serving a retrained model

Imagine a document-processing company that retrains a classifier as new document types arrive. It wants each candidate evaluated before replacing the serving version. This is a proposed design based on current documentation, not a workflow run by Sequenced. Begin with the current production artifact, a fixed evaluation dataset and a written acceptance rule for each important document class.

Choose the deployment boundary first using the platform deployment guide. In a self-managed deployment, the customer operates the data-plane infrastructure and upgrades. In managed BYOC, Union operates that infrastructure inside the customer's cloud account. Both use a Union-hosted control plane in the documented standard architectures, but the access and maintenance responsibilities differ.

Define separate task environments for preprocessing, GPU training and evaluation. Pin images and specify resources appropriate to each stage. A shared environment is convenient when tasks need the same dependencies; separate environments are useful when a CPU transformation should not inherit a large training image or request a GPU it cannot use. Record the choices so a later cost or compatibility issue can be traced to configuration.

Register the resulting dataset and model through the artifact guide. An artifact is a named, versioned reference to an object in storage, with a record of the producing run and consuming work. Keep the underlying objects available for the retention period required by the release process. A lineage link cannot reconstruct a model file that was deleted outside the platform.

Fan out evaluation over the selected document groups, then collect the results into a report comparing the candidate against the currently served model. The software-development lifecycle guide explains why data and model changes require more than ordinary code checks. Our proposed acceptance gate should examine the distribution of errors and explicitly reject a candidate that damages a critical document class despite improving the aggregate.

Test interruption in a controlled environment. Stop one evaluation worker after other partitions finish, then inspect how the workflow recovers and what it recomputes. Separate pure computation from external side effects. A retry of a model-scoring task has different consequences from a retry that writes a customer-facing record, so the latter needs an application-level strategy for recognizing already-completed work.

Keep model publication behind an explicit decision. Once the candidate is accepted, configure the serving app to use the selected artifact version and verify that requests reach that version. Preserve the previous artifact and deployment configuration so recovery means a known rollback, not another attempt to train yesterday's model from memory. The app and task histories should make that transition understandable.

Finally, reconcile the workflow's task structure with its resource usage. Very small tasks can improve visibility but also add orchestration and billing activity. Very large tasks can be harder to recover without repeating expensive work. Measure a representative workload before deciding where to place boundaries, and include queueing and startup time in the operational result.

04 / PricingA monthly minimum covers usage, while managed deployment is additional

The pricing page, consulted 7 October 2026, lists a 30-day self-service trial followed by a $950 Team monthly minimum credited against usage. Usage is metered across actions, allocated vCPU, memory and GPU. The first action band is $0.003 per action, with marginal volume bands that reset monthly. A cache hit still counts as an action; retries do not create additional action counts.

The optional managed BYOC fee is $2,500 per environment per month outside the plan minimum. Self-managed infrastructure has no such management fee. Customer cloud resources remain a separate cost. Estimate the full workload using requested resources and time, not only successful model outputs, and confirm marketplace billing and any Enterprise commitments before running a sustained pilot.

RouteCommercial basisDecision boundary
Team$950/month minimum credited to usageBegins after the 30-day self-service trial
UsageActions plus allocated CPU, memory and GPUUse the current marginal rate bands
Managed BYOCOptional $2,500/environment/monthAdditional to the plan minimum and cloud costs
EnterpriseCustom commitmentsConfirm control-plane, support and scale requirements

Selected Union.ai pricing, consulted 7 October 2026. Dollar amounts as displayed; customer cloud infrastructure is a separate budget item.

05 / DistinctionsArtifact identity can connect development and serving

Union.ai's useful architectural connection is that a produced model, its lineage and the service consuming it can live in the same operating vocabulary. That makes the release question more precise: which artifact version was approved, which evaluation supported it and which service currently uses it? Those are practical questions when code, data and model weights change on different schedules.

The environment-based task model also lets the team expose a supported execution path to application developers. Infrastructure specialists can define the available resources and images while developers focus on the workflow logic. This does not remove the need to understand containers or cloud permissions; it gives those decisions a defined place to live and a way to be reused.

The company overview emphasizes running in the customer's cloud. That is a meaningful deployment choice, but its value depends on the actual boundary. Teams should evaluate the data-plane design, operations and support access together. A cloud-account label alone does not answer who can upgrade the cluster or investigate a failed deployment.

06 / QuestionsValidate recovery semantics and the retained evidence

Inspect how the installed Flyte version treats caching, task retries and external conditions for the actual workflow. The documentation spans versioned v1 and v2 material, and this blueprint follows the current v2 guides. Do not combine a v1 code example with v2 assumptions about task environments merely because both are under the same documentation domain.

Check what the chosen plan retains and where the team will keep longer-lived release evidence. Training records can remain important after short operational log windows have passed. Preserve the artifacts, evaluation report and approval record according to the model's lifecycle, with storage policies that agree with the retention expectation.

This public-source review did not run a Union cluster, test its security architecture or benchmark recovery speed. A useful proof includes a deliberately interrupted job, a controlled external side effect and a deployment rollback. The team should be able to explain the final state from the evidence even if the person who launched the run is unavailable.

07 / DecisionChoose the runtime when failures have become a workflow problem

Union.ai is worth considering when independent AI jobs have become an interdependent production process. Begin with one expensive evaluation or retraining path and make its outputs, recovery behavior and resource bill observable. Expand after the team can trace an approved artifact into serving and recover a failed run without guessing which work was already done.

ML platform team

Gate a retrained model

Connect a fixed evaluation set, candidate artifact and explicit acceptance decision before changing serving.

Prove release traceability
Agent engineering team

Exercise a failed branch

Interrupt a representative workflow and verify recovery without repeating an external write.

Define the side-effect boundary
Infrastructure owner

Choose who operates the data plane

Compare self-managed responsibilities with the optional managed BYOC fee and actual support access.

Make ownership concrete
What should we explore next?

A business worth understanding.

Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.

Suggestions are free. Selection and publication stay with the desk.

Sources

Continue reading

All in this category