About this role
<p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">We’re hiring an Applied AI Engineer to own the AI stack end-to-end — from the cloud infrastructure it runs on, through the model layer, to the systems and evaluations that keep it working in production. This is a builder role that flexes with the situation: some weeks you’re heads-down shipping as an individual contributor, others you’re setting architectural direction and leading a small team. We want someone who can do both well and read which the moment calls for.</span></p> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">You’ll work across the full model landscape. That means getting the most out of frontier APIs (Claude, GPT, and their peers) and deploying, fine-tuning, and operating open-source models when that’s the right call — on cost, latency, control, or privacy. Knowing which to reach for, and why, is a core part of the job.</span></p> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">We’re looking for someone who can apply first principles about why a transformer produces the output it does, drive real results out of frontier models, run open source models in production, and design the cloud infrastructure to serve them reliably and cost-effectively at scale.</span></p> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt"><strong>The opportunity:</strong></span></p> <ul style="margin-top:0.0cm;margin-bottom:0.0cm;text-align:justify" type="disc"> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Own the architecture of production AI systems — inference stacks, fine-tuning pipelines, retrieval and evaluation infrastructure, monitoring </span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Build on frontier models (Claude, GPT, and peers) with real rigor: tool use, structured outputs, context and cost management, evals, and guardrails — not just prompt-and pray</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Deploy and operate open-source models (Llama, Qwen, Mistral, DeepSeek, and whatever comes next) on our cloud environment — including quantization, serving frameworks (vLLM, TGI, SGLang, TensorRT-LLM), and multi-GPU inference.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Make the frontier-vs-open-source call deliberately, on cost, latency, control, and data sensitivity grounds — and be able to defend it</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Design the cloud infrastructure underneath it all: GPU orchestration, autoscaling, cost controls, VPC/networking, IAM, observability. This is not a “hand it to DevOps” role</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Fine-tune, distill, and evaluate models against real task metrics — not vibes, not leaderboards.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Track the current research literature (arXiv, major labs, key conferences) and make judgment calls on what’s ready for production versus what’s still noise</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Depending on the project, contribute independently as a senior IC or lead and mentor a small team — and switch between the two as needed</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Partner with product and leadership to translate ambiguous problems into systems that actually ship</span></li> </ul> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt"><strong>To qualify for the role, you must have:</strong></span></p> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="text-decoration:underline;font-family:arial, helvetica, sans-serif;font-size:12.0pt"><strong>Cloud & infrastructure engineering</strong></span></p> <ul style="margin-top:0.0cm;margin-bottom:0.0cm;text-align:justify" type="disc"> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">6+ years of software / infrastructure engineering, with deep production experience on at least one major cloud (AWS, GCP, or Azure).</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Strong command of GPU infrastructure: instance selection, driver / CUDA stack, containerization, Kubernetes or an equivalent orchestrator, autoscaling patterns for inference workloads.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">IaC discipline (Terraform / Pulumi / CDK), CI/CD, monitoring (Prometheus / Grafana / OpenTelemetry), and cost management as second nature.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Fluent in Python; comfortable in at least one systems-adjacent language (Go, Rust, C++) for the parts that need it. </span></li> </ul> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt"><strong><span style="text-decoration:underline">ML / AI depth</span><s> </s></strong></span></p> <ul style="margin-top:0.0cm;margin-bottom:0.0cm;text-align:justify" type="disc"> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Genuine understanding of transformer internals: attention (including variants like MQA, GQA, sliding-window, and flash attention), positional encodings (RoPE, ALiBi), tokenization, KV cache mechanics, sampling, and where each part contributes to latency, memory, and quality</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Proven ability to get production-grade results from frontier models — Claude, GPT, and peers — including tool use / function calling, structured outputs, retrieval, context and cost management, and building the evals and guardrails around them.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Hands-on experience fine-tuning open-source LLMs — full fine-tuning, LoRA / QLoRA, preference optimization (DPO / ORPO / equivalents) — and knowing when each is appropriate.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Practical familiarity with the modern training and inference stack: PyTorch, Hugging Face, DeepSpeed or FSDP, vLLM or equivalent, evaluation frameworks</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Ability to read a recent paper, explain what’s actually new, and give an honest assessment of whether it’s worth integrating.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Track record of taking models — frontier or open weights — into a production system that real users depend on, with the evaluation, guardrails, and monitoring to match.</span></li> </ul> <p style="margin:0.0cm 0.0cm 8.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt"><strong>What we look for: </strong></span></p> <ul style="margin-top:0.0cm;margin-bottom:0.0cm;text-align:justify" type="disc"> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Strong understanding of the mechanics before reaching for an abstraction.</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Able to contribute as an individual contributor, adapting to the needs of the situation</span></li> <li style="margin-top:0.0cm;margin-right:0.0cm;margin-bottom:8.0pt;line-height:115%;font-size:12.0pt;font-family:arial, helvetica, sans-serif"><span style="font-family:arial, helvetica, sans-serif;font-size:12.0pt">Communicate tradeoffs clearly to non-technical stakeholders </span></li> </ul> <p style="margin:0.0cm 0.0cm 8.0pt 36.0pt;line-height:115%;font-size:12.0pt;font-family:Aptos, sans-serif;text-align:justify"> </p>