About this role
Research Scientist, Reinforcement Learning at Deeproute.ai. Location: Fremont, California, United States. Role: training policies, scaling training, optimizing rewards Requirements: Proficiency in modern RL and RLHF algorithms, experience with reward model training and LLM/VLM fine-tuning, distributed RL training and massively parallel simulation, sim-to-real transfer, Python and C++, PyTorch, and distributed training frameworks. Category: Research and Development (R&D) Seniority: Senior Level Tools: DQN, PPO, SAC, TD3, DPO, GRPO, LLM, VLM, VLA, Python, C++, PyTorch, Ray, Horovod, CUDA Commitment: Full Time Workplace: Onsite Languages: English