Description:
As Domain Architect - AI Networking, you will be the primary technical authority for the physical and logical lifecycle of high-performance interconnect fabrics across a diverse portfolio of client environments, bridging the gap between architectural design and hands-on execution. You are a builder as much as an advisor: as comfortable configuring adaptive routing on a leaf-spine fabric from the CLI as you are explaining that configuration to a C-level audience.
As a global systems integrator, we don't simply operate static cloud environments. We design and deliver purpose-built, high-scale AI factories for some of the world's leading enterprises. In this role you will define the reference standard for network infrastructure, moving beyond single-switch administration to architect repeatable, scalable, and automated lossless fabrics. You will act as technical lead on NVIDIA Cloud Partner (NCP) and private enterprise AI cloud deployments, owning the network layer of the compute, network, and storage stack.
Your time will be split roughly 60/40 between delivering complex AI infrastructure (60%) and providing pre-sales subject matter expertise (40%). You will lead the physical provisioning of InfiniBand and Spectrum-X Ethernet fabrics for NVIDIA DGX SuperPOD, NVIDIA DGX BasePOD, and Cisco AI POD environments, ensuring clients inherit platforms that are genuinely ready for day-2 operations, while helping the sales team scope and cost future deployments.
Key Responsibilities
Delivery and implementation
High-performance fabric design
Architect and deploy non-blocking fat-tree (Clos) topologies using NVIDIA Quantum-2 (NDR) and Quantum-3 (XDR) InfiniBand switches, and NVIDIA Spectrum-4 Ethernet switches.
Implement rail-optimized network designs so GPU-to-GPU traffic aligns with compute-node PCIe topology, minimizing latency for NCCL collective operations.
Configure adaptive routing, congestion control, and quality of service (QoS) to prevent head-of-line blocking and guarantee lossless data delivery.
In-network computing and offload strategy
InfiniBand: enable and tune SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) on Quantum switches to offload collective operations (AllReduce, ReduceScatter) from the GPU to the switch silicon.
Ethernet: architect high-performance Ethernet AI fabrics using SmartNIC/DPU offload (NVIDIA BlueField-3) to accelerate collective operations and isolate management traffic from the data path.
Evaluate and roadmap emerging Ultra Ethernet Consortium (UEC) standards as the practice transitions in-network collectives from proprietary InfiniBand toward open Ethernet.
Tune RoCEv2, PFC, and ECN on Ethernet fabrics and SmartNICs to approximate InfiniBand-class lossless behavior.
Fabric management, automation, and telemetry
Deploy NVIDIA UFM (Unified Fabric Manager) to manage InfiniBand subnets and NetQ for Ethernet fabric telemetry; implement PTP for nanosecond-level clock synchronization across the cluster.
Use automation and AI-assisted tooling to generate switch configuration from the P2P cabling schedule, LLD, and PDG, then validate every topology in NVIDIA DSX Air before it ever touches hardware.
Monitor fabric health continuously to identify slow receivers and link degradation in real time, and maintain Python and Ansible-based infrastructure-as-code for ongoing configuration management.
Layer 1 precision and performance engineering
Own the physical cabling strategy, defining cable schedules (DAC, AOC, or OSFP transceivers) that meet signal-integrity requirements over the required distances.
Validate the link budget to ensure optical loss stays within acceptable limits for 400G and 800G links.
Conduct acceptance and validation testing using NCCL tests to verify fabric performance against expected baselines.
Pre-sales SME and consulting
Technical scoping and estimation
Calculate the bisectional bandwidth required for a client's workload (for example, training typically requires 1:1 non-blocking; inference may tolerate 3:1 oversubscription), and produce accurate level-of-effort (LOE) estimates for statements of work.
Bill of materials validation and architecture
Own the technical accuracy of the network bill of materials, validating that every switch, transceiver, and cable is on the NVIDIA or OEM hardware compatibility list (HCL).
Manage the complexity of breakout cables (for example, 800G to 2x400G) and connector types (OSFP, QSFP112) to prevent on-site installation failures.
Client workshops
Educate clients on the difference between standard enterprise Ethernet (lossy, TCP-based) and AI fabric (lossless, RDMA-based) traffic.
Design the integration point between the high-speed “back-end” AI fabric and the client's existing “front-end” management network, including BGP/EVPN handoffs.
Minimum Qualifications
10+ years in networking, HPC, or data center infrastructure engineering, including 7+ years in a customer-facing architecture role.
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Deep working knowledge of NVIDIA networking platforms, including Quantum InfiniBand switch architecture (MLNX-OS, UFM) and Spectrum-X Ethernet running NVIDIA Cumulus Linux or SONiC.
Proficiency in RoCEv2 (RDMA over Converged Ethernet), including PFC and ECN, and a working architectural understanding of NVIDIA BlueField DPU and DOCA offload use cases.
Automation and infrastructure-as-code experience, with proficiency in Python and Ansible for switch configuration management.
Demonstrated ability to translate technical architecture into commercial outcomes for senior stakeholders.
Preferred Qualifications
Industry background: experience within a systems integrator (SI) or managed service provider (MSP) environment.
Multi-vendor exposure: Arista EOS for high-performance AI Ethernet, or Cisco Nexus Dashboard for AI infrastructure.
Routing protocols: solid understanding of BGP, EVPN, and VXLAN for multi-tenant isolation.
Compute integration: understanding of how the network interacts with the host OS (IPoIB, Netlink) and optimization of GPUDirect RDMA (GDR) and NCCL communication patterns.
Scale-across: experience extending an InfiniBand fabric across multiple data center halls (NVIDIA MetroX).
Certifications: NVIDIA-Certified Professional: AI Networking (NCP-AIN); NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO); Arista ACE (Cloud Engineer); Cisco CCNP/CCIE Data Center.
| Organization | World Wide Technology |
| Industry | IT / Telecom / Software Jobs |
| Occupational Category | Domain Architect |
| Job Location | New York,USA |
| Shift Type | Morning |
| Job Type | Full Time |
| Gender | No Preference |
| Career Level | Intermediate |
| Experience | 2 Years |
| Posted at | 2026-09-29 3:48 pm |
| Expires on | 2026-11-13 |