About this role
Deep Learning Performance Architect at NVIDIA. Location: Shanghai or Beijing. Role: developing kernels, optimizing performance, collaborating teams Requirements: Masters/PhD or equivalent in CE/CS/AI, 2+ years relevant experience, strong C/C++ skills, GPU programming (CUDA/OpenCL), performance optimization and profiling, EMR not mentioned, Python a plus. Category: Software Development Seniority: Entry Level Tools: C, C++, Python, CUDA, OpenCL, Tensor-RT Commitment: Full Time Workplace: Onsite Languages: English