About this role
Deep Learning Performance Architect at Nvidia. Location: Shanghai or Beijing. Role: developing kernels, optimizing performance, profiling code Requirements: Masters/PhD or equivalent experience, 2 years relevant experience, strong C/C++ skills, GPU programming (CUDA/OpenCL), performance modeling/profiling, Python a plus, excellent communication. Category: Software Development Seniority: Entry Level Tools: Tensor-RT, C, C++, Python, CUDA, OpenCL Commitment: Full Time Workplace: Onsite Languages: English