About this role
Member of Technical Staff, Inference at Inferact. Location: San Francisco or United States. Role: optimizing inference, implementing architectures, debugging code Requirements: Bachelor's or equivalent, deep knowledge of transformer architectures, strong Python and PyTorch skills, experience with vLLM/TensorRT-LLM/SGLang/TGI, ability to implement research papers and produce performant, maintainable ML systems. Category: Software Development Seniority: Senior Level Tools: Python, PyTorch, vLLM, TensorRT-LLM, SGLang, TGI, verl, OpenRLHF, Unsloth, LlamaFactory Commitment: Full Time Workplace: Onsite Languages: English