About this role
LLM Engineer (Optimization) at 42dot. Location: Pangyo, Gyeonggi-do, Korea. Role: optimizing inference, developing runtime, analyzing performance Requirements: 3+ years in LLM/ML infrastructure or inference optimization; experience building inference engines/runtimes; knowledge of GPU architecture, CUDA, quantization and model compression; proficient in Python or C/C++ and deep learning frameworks. Category: Engineering Seniority: Mid Level Tools: vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX, CUDA, Triton, TensorRT, TVM, MLIR, NVIDIA Nsight Systems, Nsight Compute, py-spy, PyTorch, ONNX, Python, C/C++ Commitment: Full Time Workplace: Onsite Languages: Korean