About this role
* Assess AI training and inference workloads and translate business and technical requirements into
infrastructure architecture, system specifications and capacity plans.
* Evaluate heterogeneous computing platforms, including Intel and AMD CPUs, NVIDIA GPUs, and
accelerators such as Ascend NPUs and Google TPUs, to determine workload suitability and identify
potential performance bottlenecks.
* Develop plans for computing resource pooling, infrastructure expansion and long-term capacity
requirements, considering performance, scalability, compatibility and power consumption.
* Design and integrate GPU servers, high-performance storage, RDMA networks, distributed training
frameworks, and containerization and virtualization platforms into coordinated infrastructure
solutions.
* Research, design and evaluate network architectures and communication topologies to support
efficient data transfer and distributed computing across cluster nodes.
* Conduct system integration, compatibility and performance testing; document findings and
recommend improvements to infrastructure configurations.
* Optimize CPU and accelerator utilization, memory bandwidth, storage throughput and inter-node
communication to improve training throughput, reduce inference latency and increase overall
resource efficiency.
* Oversee the technical operation of IT infrastructure, monitor system availability and performance,
investigate bottlenecks and coordinate corrective actions.
* Lead and coordinate technical personnel and service providers in the implementation, integration and
optimization of infrastructure solutions.
* Provide advanced technical advice to clients and internal teams on infrastructure selection,
deployment, upgrades and performance improvement.
* Prepare and maintain architecture diagrams, configuration specifications, capacity assessments, test
reports and operational documentation.
Job Type: Full-time
Pay: $65.00-$68.00 per hour
Work Location: In person