About this role
Our client
WebPros, the largest web hosting software and automation company, manages 900,000+ servers, 85 million domains, and 33 million users. WebPros unites top providers in web hosting, billing automation, infrastructure, server management, and online marketing software. Currently, their lineup includes cPanel, Plesk, SolusVM, WHMCS, XOVI NOW, Sitejet, 360 Monitoring, and koality, with ongoing additions.
About the role
Join our global Data & AI team at WebPros, where you will operationalise and optimise
Large Language Model (LLM) solutions that power our products. You will work across
product and engineering teams to design and evaluate prompts and agentic workflows,
put safety guardrails in place, and keep our generative AI systems observable, traceable
and cost-efficient.
This is a hands-on, technical role with a high degree of autonomy. We are looking for
someone who can take an abstract goal, break it down into concrete steps, and drive it
to completion.
How you'll work (AI-native)
We work AI-native, and we expect the same from this role:
• You develop and explore code primarily through AI coding agents (e.g. Claude
Code), including navigating and changing large, unfamiliar codebases.
• You practise agentic engineering: a structured engineering process where intent
is the source of truth, and where you rigorously review what the agents produce
to confirm it is correct and well-designed.
• You are comfortable with the common building blocks of harness engineering
(e.g. MCP, Skills and Plugins, subagents, spec-driven development frameworks
such as Spec Kit) and know when to reach for each.
• We don't require extensive prior experience in software development, but you
need to comfortably reach correct, well-reasoned results with AI agents.
Therefore, experience programming in popular scripting languages (e.g.
TypeScript, Python) is required.
Responsibilities
As an AI Engineer, your key responsibilities will include:
• Designing, testing and optimising agentic design patterns using techniques such
as zero-shot, few-shot, chain-of-thought, ReAct, and others.
• Designing and comparing agentic architectures (single-agent loops, multi-agent
systems, prompt-template pipelines) and choosing the right one for the task on
cost, latency and quality.
• Evaluating and recommending models, both small and large, across closed-
source, and open-weights options, and designing model routing and selection to
balance quality, cost and latency.
• Building and running strong evaluation pipelines, offline and online: assembling
datasets, defining metrics (similarity scores, LLM-as-a-judge, cost, latency,
quality) and tracking generation quality over time.
• Implementing guardrails through input/output validation, including prompt-
injection prevention, toxicity filtering and hallucination detection.
• Integrating our systems with LLM observability platforms to monitor latency,
token usage, error rates and cost, and using that visibility to inform guardrails
and optimisation.
• Integrating LLM workflows into our products and services.
• Supporting AI risk assessments and compliance efforts.
What we're looking for (required)
To be successful in this role, you should bring:
• An AI-native working style: fluency with AI coding agents (e.g. Claude Code) for
building and exploring code, and the judgement to assess what they produce.
• Autonomy: a track record of taking an abstract goal, scoping it yourself and
delivering it with little supervision.
• Hands-on prompt engineering and optimisation, including agentic prompting
(experience with an optimisation framework such as DSPy is welcome).
• Practical experience with testing, evaluating, and red teaming LLM applications:
designing evals, building datasets, defining testing strategies, and comparing
and selecting models (frontier, and open-weights), with tools like Promptfoo.
• Experience with LLM observability, monitoring, logging and traceability (e.g.
Langfuse, MLflow, etc.).
• Familiarity with AI guardrailing techniques (e.g. Guardrails AI, LLM Guard, ...).
• Comfort working with LLM APIs and SDKs in code (e.g. Anthropic Claude,
OpenRouter, Pydantic AI, LangChain, Vercel AI SDK, etc.).
• An understanding of AI governance and risk management practices.
• Strong English (C1 level), with the ability to explain complex AI concepts clearly
to both technical and non-technical audiences.
Nice to have
• Experience with open-weights models and fine-tuning, ideally building a
reusable training pipeline (data, pipeline, model) that can be pointed at new
tasks rather than a one-off.
• Experience with text-generation use cases, which is where most of our work sits.
• Experience setting up and working with MCP and A2A.
_____
Q & A:
— Does the job come with a probation period, and if so, how long does it last? Yes, there is a 3-month probation period. — What is the expected work schedule? Full-time, flexible. You can work remotely and also you can choose hybrid mode where you can combine working on-site (in Lviv office) and remotely. — How many vacation and sick days are provided? Annual paid vacation — 20 working days/ 7 unconfirmed sick days/days off a year.
Social package & benefits:
• Full medical insurance
• MacBook & accessories
• English lessons
• Accountant assistance
• Minimal bureaucracy, synergy, and formalities, primarily focusing on effective communication
•
Hiring process:
• Screening call with Recruiter (soft skills interview) ~ 20 min
• Intro call with the Team Lead and team members ~ 30-45 min
• Technical interview ~ 1,5 hour
Join Our Team!