Demo — AI Data Center Fabric Test Methodology

Videos

Optimizing AI Fabrics: Keysight's Automated Test Methodology for AI Data Centers

In this session, Alex Bortok, Lead Product Manager for AI Data Center Solutions at Keysight Technologies, provides a comprehensive overview of Keysight’s AI fabric test methodology. This approach is designed to guide engineers and data center architects through each phase of AI fabric design, validation, and optimization, with a focus on achieving high-performance, low-latency, and balanced network behavior.

 

Through automated testing, parameter optimization, and real-world demonstrations, Alex highlights how data center teams can enhance the reliability and efficiency of their AI training and inference fabrics, ensuring scalability and top-tier performance for AI workloads.

 

Understanding AI Fabric Design and Test Phases

Keysight’s AI fabric methodology supports a systematic process for designing, testing, and fine-tuning AI backend networks. These backend "fabrics" are critical to the performance of distributed AI workloads, especially during collective operations such as broadcast, all-reduce, and all-to-all exchanges across GPU clusters.

 

The methodology emphasizes the following design elements:

  • Topology selection (e.g., fat-tree, dragonfly, or custom mesh)
  • Collective operation algorithms
  • Performance isolation
  • Load balancing strategies
  • Congestion control mechanisms

By simulating and analyzing these factors in a controlled environment, engineers can predict real-world performance outcomes and avoid costly deployment missteps.

 

Metrics That Matter: Evaluating AI Fabric Performance

The performance of an AI fabric can be evaluated using several critical metrics. Alex introduces the following:

  1. Collective Completion Time
    The total time it takes for a collective operation (e.g., all-reduce) to complete across all ranks.

  2. Algorithm Bandwidth
    The effective data transfer rate achieved by the collective algorithm across the network.

  3. Bus Bandwidth
    A normalized performance metric that accounts for underlying system limits, such as PCIe or memory bandwidth. Bus bandwidth is especially valuable because it removes the number of GPUs from the equation, isolating the true limiting factor of the operation.

By focusing on bus bandwidth, engineers can better understand whether a slowdown is caused by network constraints, memory access, or compute bottlenecks.

 

Testbed Setup: Simulating Real AI Cluster Conditions

To demonstrate the power of this methodology, Keysight built a realistic testbed that simulates the backend environment of an AI data center. The setup includes:

  • Four switches with 800 Gbps port speeds
  • Emulation of 16 GPUs or NICs running at 400 Gbps each
  • Synthetic AI traffic and collective communication patterns
  • Automated tools for real-time performance measurement and tuning
  • This testbed allows for the emulation of production-scale GPU clusters, enabling precise testing without the need for physical GPUs or full-scale infrastructure.

 

Congestion Control and Bandwidth Optimization

One of the session's most impactful demonstrations is the comparison of fabric performance with and without congestion control enabled.

 

Key Takeaways from the Congestion Control Test:

Without congestion control: Network bandwidth is underutilized, and packet loss or retransmissions occur.

  • With congestion control (DCQCN) enabled: Bandwidth utilization is significantly improved, and congestion is minimized.
  • DCQCN (Data Center Quantized Congestion Notification) is a key protocol in modern data centers for managing congestion on high-speed fabrics. Fine-tuning DCQCN parameters—such as alpha, target rate, and update interval—can lead to substantial performance gains.
  • By leveraging Keysight’s automated tools, engineers can test multiple parameter combinations, visualize the results, and converge on the optimal fabric configuration that balances performance, reliability, and fairness.

 

Benefits of Automated AI Fabric Testing

Keysight’s AI fabric testing methodology provides the following advantages for data center teams:

  • Faster validation cycles through automation
  • Reduced hardware costs by emulating GPUs and NICs
  • Repeatable, scalable testing of complex topologies
  • Actionable performance data for design decisions
  • Early detection of bottlenecks and scalability issues

This method empowers organizations to design AI clusters that are production-ready, even before physical deployment begins.

 

Industry Collaboration: Ultra Ethernet Consortium

As part of its commitment to innovation in AI networking, Keysight is an active member of the Ultra Ethernet Consortium—a group dedicated to developing next-generation Ethernet-based solutions tailored for AI and HPC workloads.

Keysight continues to collaborate with other industry leaders to advance standardized performance benchmarks, interoperability testing, and future-proof data center architectures.