Keysight Technologies helps engineers test the infrastructure that AI systems depend on. Its AI Data Center Builder focuses on distributed communication: how collective operations and workload traffic interact with a network fabric. The offer is useful when a team needs to understand bottlenecks before committing to a larger cluster or a new switching design.
- 01The offer AI workload emulation, collective benchmarking and network measurement.
- 02The fit Infrastructure operators and equipment teams validating AI cluster fabrics.
- 03The boundary Public documentation and a proposed comparison experiment; no network, accelerator or Keysight trial was run.
01 / ProductEmulate the communication workload instead of assuming the network is ready
AI Data Center Builder brings workload emulation and collective benchmarks into a configurable test environment. Its product page lists RDMA and RoCEv2 support, common collective operations and hardware or software configurations. The target is network behavior associated with AI workloads, not the semantic quality of a language model.
The documentation overview describes several test-engine choices: RoCEv2 endpoint emulation on AresONE hardware, software endpoints on servers with RDMA network adapters, and real AI accelerators. Those choices change what is physically present in an experiment and which assumptions need to be validated.
The architecture guide separates a Distributed System Experiments controller, Data Flow Emulation components and persistent storage for trial configurations and results. That separation is useful for repeatability. A benchmark is more informative when its traffic definition and environment are retained, rather than reduced to one untraceable headline number.
02 / AudienceFor the team deciding how a cluster should move data
A strong user is an infrastructure team considering a fabric upgrade before expanding its accelerator fleet. It wants to know how routing, congestion controls and competing jobs affect communication completion. Another reader is an equipment vendor that needs reproducible evidence for a switch or adapter configuration.
This is not a substitute for evaluating model accuracy, application correctness or the complete economics of an AI service. Emulating network traffic can isolate a bottleneck without executing all the computation that created the original traffic. The resulting evidence has a defined scope and should be carried into a later application-level test.
Arista’s AI networking approach is relevant to the fabric being selected, while NVIDIA’s infrastructure stack offers context for accelerator and communication choices. Keysight occupies the measurement layer across those decisions. A useful comparison asks which test distinguishes competing designs, rather than which vendor supplies the largest collection of components.
03 / WorkflowProposed workflow: compare two fabric configurations before expansion
Begin with a proposed cluster expansion that must support several concurrent training jobs. Write down the decision the experiment will inform: for example, whether a changed routing or congestion configuration improves completion behavior under contention. Keep the physical fabric and all unrelated settings fixed where possible so that the comparison has an interpretable cause.
Select representative communication patterns from the intended workload. A single bulk-throughput stream is unlikely to capture the timing of collective communication. Include message sizes and synchronization patterns relevant to the actual jobs. Record how the chosen profiles were derived, and identify where they simplify the production workload.
Choose a test-engine route according to the question and available lab resources. Software endpoints on RDMA-equipped servers may fit an existing lab; high-density traffic equipment may support a different scale or measurement requirement. The emulation method itself should be included in the result so that another engineer can understand what was exercised.
Check the deployment prerequisites before scheduling the comparison. The documentation distinguishes controller requirements, AresONE synchronization and server/NIC requirements. It specifically emphasizes precise time synchronization between software DFE hosts. Poor timing alignment can undermine a latency comparison even when traffic generation appears to work.
Establish a baseline with one job and a simple communication pattern. Verify connectivity, expected endpoint participation and the relationship between configured and observed traffic. Retain the baseline configuration before adding contention. If the simple experiment is not understood, a more complex workload mix will make diagnosis harder rather than more realistic.
Introduce multiple jobs with controlled overlap. Measure completion-time distribution and inspect whether one workload is disproportionately affected. An average throughput improvement may conceal a long tail or poor isolation for a smaller job. Choose metrics that correspond to the capacity decision, not only those that produce the most attractive dashboard.
Change one fabric setting at a time and repeat the same trials. Keep run order and background traffic visible. If a difference is small relative to run-to-run variation, the result should remain uncertain. This proposed method does not claim that any particular switch setting will improve a production cluster.
Inspect bottlenecks in the context of the communication graph. A congested link can reflect placement or traffic synchronization rather than insufficient aggregate bandwidth. Use the experiment to distinguish those explanations. That distinction may support a placement or configuration change before a larger equipment purchase is justified.
After selecting a promising configuration, compare the emulated result with a bounded real-workload test. The aim is to understand where the emulator is predictive and where compute, storage or software behavior changes the conclusion. Preserve discrepancies as evidence; do not discard them merely because the larger synthetic test looked favorable.
Package the result with the endpoint count, topology, firmware versions, traffic definition and synchronization method. The next engineer should be able to reproduce the comparison or identify why the production environment differs. A measurement report that survives configuration changes is more useful than a one-time claim that the fabric is fast.
04 / PricingLicenses cover applications and concurrent transport endpoints
| Item | Commercial or access basis | Practical implication |
|---|---|---|
| Product configuration | Quote-based hardware bundles or software options | Request the exact application, interface and scale required. |
| Collective Benchmarks and Workload Emulation | Valid application licenses required to run trials | The word trial in the guide describes an experiment, not an assumed free evaluation. |
| Universal Transport Endpoints | Quantity based on concurrent endpoints or NICs | Count the intended simultaneous experiment size. |
| License server | Separate deployment enables floating licenses | Plan controller connectivity and license activation. |
| Results viewing | Previous results can be viewed without a valid run license | Viewing evidence does not establish permission to run another experiment. |
Commercial scope from product configurations and KAI licensing documentation. Consulted 3 October 2026.
The licensing guide is more precise than a general request-a-quote button. It distinguishes application entitlements from endpoint quantities and describes a license server that allows licenses to float between instances. Hardware-emulated NICs and software use of RDMA NICs both have endpoint requirements.
This review did not establish a public universal price for the proposed configuration. A quote should include the intended concurrent endpoint count, applications, physical hardware if needed and support arrangement. A software-only option still requires suitable servers, adapters and lab capacity; it is not a statement that the whole test environment is free.
Include the effort to maintain reproducible configurations in the evaluation budget. If the team cannot recreate its baseline after a software update, the apparent savings from a quick benchmark may be lost in repeated investigation. Retained trial definitions and results are practical assets for later network changes.
05 / DistinctionsThe experiment can isolate communication from other cluster costs
Keysight’s AI networks portfolio spans network and interconnect testing. Its measurement tools help engineers evaluate the infrastructure around an AI workload. As clusters become more complex, understanding the path between compute nodes can be a separate engineering problem from choosing the nodes themselves.
Our assessment is that workload emulation is most useful when it isolates a question that a full-cluster trial would make expensive or difficult to repeat. It can support comparison and diagnosis. Its value is not that it removes every need for real application testing, but that it can make those later tests more focused.
06 / QuestionsAsk which parts of production the experiment represents
How representative is the traffic? A trace or collective profile can omit compute delays, storage behavior or application scheduling. State those omissions and test whether they could alter the decision. A synthetic workload should not be presented as a complete application result.
Are the endpoints and clocks suitable? The prerequisite page marks some server guidance as preliminary and names tested NIC families. Confirm support for the intended lab configuration and software release. “Other adapters may work” is a different claim from documented validation.
What result changes the purchase or configuration decision? Set that criterion before running a large matrix. If every possible outcome leads to the same equipment choice, the experiment may be collecting interesting data without resolving the actual uncertainty.
07 / DecisionChoose Keysight when a repeatable network question justifies the testbed
The next useful step is a small baseline and one controlled contention experiment for a real infrastructure decision. Expand the matrix only after the team understands the endpoint behavior and measurement scope. A result that explains why one configuration handles the intended workload better is more actionable than a generic throughput score.
Cluster operator
Compare one consequential fabric setting under representative concurrent jobs.
Network equipment team
Retain topology, traffic and timing details with each benchmark.
Infrastructure buyer
Quote applications and concurrent endpoints, then budget the complete lab.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- AI Data Center Builder product and configurationsConsulted
- KAI documentation overviewConsulted
- System architectureConsulted
- Deployment prerequisitesConsulted
- Licensing requirementsConsulted
- AI networks portfolioConsulted


