About this role
Member of Technical Staff, Inference at Radical Numerics. Location: San Francisco, California, United States. Role: improving performance, developing kernels, deploying infrastructure Requirements: Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment. Category: Engineering Seniority: Senior Level Tools: CUDA, Triton, Python, PyTorch, vLLM, TensorRT-LLM, SGLang, DeepSpeed Commitment: Full Time Workplace: Onsite Languages: English