About this role
The Role
Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?
That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.
This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.
What you get to do in this role:
We're hiring across three areas. You'll anchor on one and touch the others; which one is a conversation we have with you, not a slot we drop you into.
Eval orchestration at scale
The runtime that executes multi-turn agent scenarios end-to-end — stand up the environment and user simulator, drive the user↔agent↔world loop, collect transcripts, traces, and final state, run validators and scoring, tear downScheduling, retries, high-concurrency execution, and run isolation at production dataset sizesVersioned specs, datasets, and reports, with run-to-run comparison as a first-class operationConsolidating evals that run today as one-off workflows onto a single orchestration service — one source of truth, one place to schedule and retryEstablishing a reliability floor and an SLO for the harness itselfGetting to self-serve, so any team runs an eval without bespoke integrationAgent observability and tracing
Leading the move to OpenTelemetry-native observability for the agent platform, replacing the parallel per-service logging, correlation, and redaction mechanisms in use todayThe span data model for agent trajectories — prompts, tool calls, plan updates, outcomes — so a trajectory is queryable, not reconstructed by hand from log filesTrace context propagation across async boundaries and sessions that stay alive for minutes or hoursMaking full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own dataFault attribution and cross-run diffing: which component actually broke, and what changed since the last green runThe debug surface support and harness engineers use, and the tracing contract with the team that builds the agentStateful simulation
The simulation environment itself: stateful fakes of the enterprise systems agents call — ITSM, HR, knowledge bases, inventory — backed by a real datastore that persists changes during a run, so a created ticket is visible to a later readPer-run data injection and programmatic setup/teardown so every run is hermetic and repeatableLLM-driven user simulators for open-ended personas, and scripted state-machine simulators for deterministic flowsContract-testing mocks against real API schemas in CI, so simulation fidelity can't quietly drift as vendor APIs changeAhead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads — provisioned from an identical baseline and torn down every runAnd across all three: laying the foundation for using eval signal to optimize the agent, not just measure it.
To be successful in this role you have:
Experience in at least 3 of these:
Distributed systems: idempotency, delivery guarantees, isolation, and — unusually central here — determinism and reproducibilityOrchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high-concurrency job systems (Temporal, Airflow, Argo, or something you built yourself)Observability internals as a builder, not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace dataConcurrent and async programming: Python asyncio, Go concurrency, structured cancellationData-intensive pipelines: high-volume ingest, schema evolution, sampling and retention trade-offsgRPC/protobuf service and interface designRequired:
8+ years building production backend or infrastructure systemsStrong in Python or Go (ideally both)Experience designing and operating systems that handle real traffic at scaleComfort making a non-deterministic system measurable. You don't need an ML background — but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases onComfort with ambiguity; these are novel problems without textbook solutions
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license.