About this role
What You’ll Do As a Senior DevOps Engineer, you’ll own the design, reliability, and evolution of InteractiveAI’s production infrastructure, ensuring our platform is secure, scalable, observable, and fast to ship. You’ll guide a small, high-impact team of DevOps/SRE engineers, raising standards for uptime, deployment velocity, and operational excellence across cloud and hybrid environments.<br>You’ll work closely with AI Engineers, Backend Engineers, and Product teams to support demanding workloads (including LLM/ Agent orchestration) while building a best-in-class AI engineering platform.<br>Your responsibilities will include:<ul><li>Architect and evolve InteractiveAI’s core infrastructure (networking, compute, storage, Kubernetes, CI/CD) to support rapid product delivery</li><li>Build and operate cloud-agnostic environments (Kubernetes) across VPC, on-prem, and hybrid deployments</li><li>Implement and maintain CI/CD for services and infrastructure (GitOps preferred), enabling safe and frequent deployments</li><li>Define “gold standards” for reliability: SLOs/SLIs, incident response, on-call practices, runbooks, and postmortems</li><li>Establish robust observability (metrics, logs, traces) with actionable alerting and performance monitoring</li><li>Drive infrastructure-as-code across environments (Terraform/Pulumi/CloudFormation), including strong review/testing practices</li><li>Own security fundamentals: secrets management, IAM, network segmentation, vulnerability management, and compliance readiness</li><li>Improve developer productivity with internal tooling, self-service environments, and paved paths</li><li>Optimize costs and capacity planning across clusters and workloads (including bursty inference/agent workloads)</li><li>Mentor and grow a small DevOps team, fostering ownership, quality, and pragmatic execution</li></ul> What We’re Looking For We’re seeking a hands-on technical leader with a builder’s mindset—someone who thrives in fast-moving environments, can scale systems and practices, and can lead through influence and example.<br>Minimum Requirements<ul><li>5+ years in DevOps, SRE, Platform Engineering, or Infrastructure roles</li><li>Proven track record operating production systems with strong reliability and incident management practices</li><li>Strong expertise with Linux, networking fundamentals, and cloud infrastructure patterns</li><li>Deep knowledge of containerization and orchestration (Docker, Kubernetes)</li><li>Experience building and owning CI/CD pipelines for microservices and infrastructure</li><li>Proficiency with infrastructure-as-code (Terraform, Pulumi, CloudFormation)</li><li>Strong scripting/coding ability (Python, Bash, Go, or similar)</li><li>Experience with observability stacks (Prometheus/Grafana, ELK/OpenSearch, Datadog, OpenTelemetry, etc.)</li><li>Experience with at least one major cloud provider (AWS, GCP, or Azure)</li></ul>Additional Requirements (Nice to Have)<ul><li>Experience with service mesh, API gateways, or ingress at scale (Istio/Linkerd, NGINX, Envoy, etc.)</li><li>Familiarity with security & compliance practices (GDPR, ISO 27001, SOC 2 readiness)</li><li>Experience supporting GPU clusters or high-throughput compute workloads (a plus for AI platforms)</li><li>Exposure to MLOps/LLM serving infrastructure and performance/cost optimization</li><li>Prior experience leading or mentoring engineers (formally or informally)</li></ul> What You’ll Get <ul><li>Competitive base salary (€90,000/yr to €110,000/yr) + performance bonuses</li><li>Access to equity/share plan as it rolls out</li><li>Private health insurance</li><li>Flexible work setup + travel when needed (hybrid in Lisbon or Madrid)</li><li>23 days PTO (excluding local public holidays)</li></ul> Who You Are <ul><li>Proactive & Strategic: You anticipate scaling challenges and design systems that won’t collapse under growth.</li><li>Technical Leader: You raise the bar with strong engineering fundamentals, clarity, and high standards.</li><li>Accountable & High-Ownership: You take responsibility for uptime, security, and delivery outcomes.</li><li>Builder Mentality: You move fast in ambiguity while keeping production safe and predictable.</li><li>Collaborative Partner: You communicate clearly, build trust across teams, and balance pragmatism with long-term vision.</li></ul> Interview Process We keep our process focused and respectful of your time. Most candidates complete it in 2–3 weeks.<br>Here’s what to expect:<ol><li>Intro Call – 30 minutes with our team to align on fit and expectations</li><li>Technical Challenge – A practical DevOps/SRE design or automation task</li><li>Technical Interview – Deep dive into infrastructure architecture, reliability, security, and delivery practices</li><li>Leadership & Values Interview – Assess alignment with InteractiveAI’s culture and growth mindset</li><li>Offer – Final conversation and offer</li></ol>We’re building a team of builders — people who care about impact, quality, and growth.<br>If that’s you, let’s talk !