About this role
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
About Lilly At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you're driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly. About Tech@Lilly At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Bengaluru builds the capabilities that make this possible — cloud platforms, AI systems, and automation at enterprise scale — all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve. About the Organization The Clinical & Non-Clinical Data Organization at Eli Lilly and Company is responsible for the design, build, and operation of enterprise data platforms that power drug discovery, clinical development, and regulatory submissions. Data Hub is building a robust Data Strategy to make Lilly's Clinical and Non-Clinical data AI-ready and audit-ready, delivering scalable, governed, and reusable data products that accelerate how medicines reach patients. The data engineering organization sits at the intersection of science, technology, and patient impact — connecting Clinical and Non-Clinical data across the full chain, from ingestion to consumption. Path/Level: R5 (Senior Data Architect) Position Summary The Senior Data Architect (R5) is a hands-on leader who designs and personally builds AI and agentic solutions that operate at enterprise scale across Lilly's Clinical and Non-Clinical data domain — multi-agent systems, LLM-orchestrated pipelines, and retrieval/reasoning architectures built on governed, semantic data foundations. This is a builder role: the architect writes code, stands up agent frameworks, and ships production AI systems personally, not just specifications. This role is split 70% hands-on technical execution including coding and 30% strategy, and shapes the future technology landscape. A defining mandate of this role is Right Model, Right Task — routing every agentic and LLM workload to the model best suited to it on cost, latency, and accuracy grounds, integrated directly with Lilly's internal data platform so routing and reasoning are grounded in enterprise-native lineage and knowledge graphs rather than bespoke, disconnected metadata. Agentic and AI-assisted ways of working are the expected default across every design, analysis, and documentation activity — and this leader is the reference point for how the broader India team scales AI-native architecture. Key Responsibilities AI & Agentic Solution Architecture (Hands-On)
• Architect and personally build multi-agent and LLM-orchestrated solutions — planning/tool-calling agents, retrieval-augmented generation (RAG), and agent-to-agent workflows — for Clinical and Non-Clinical use cases.
• Design for scale from the start: agent orchestration, state management, concurrency, cost/latency budgets, and failure/retry handling across high-volume production workloads.
• Implement evaluation, guardrails, observability, and human-in-the-loop patterns so agentic systems are safe, auditable, and production-ready in a regulated environment.
Right Model, Right Task – Platform-Integrated Routing, Lineage & Knowledge Graphs
• Design and implement a right-model-right-task routing layer that selects the optimal model — by size, provider, and fine-tuned vs. general-purpose — for each agentic task based on complexity, cost, latency, and accuracy requirements, rather than defaulting to one model for every job.
• Integrate directly with Lilly's internal data platform (Data Hub catalog, lineage, and metadata services) rather than building parallel metadata stores, so routing and agent reasoning are grounded in the enterprise's single source of truth.
• Build and maintain lineage-aware enterprise knowledge graphs sourced from the internal data platform, capturing data provenance, sensitivity, and domain context that ground both agentic reasoning and model-selection decisions.
• Use lineage and knowledge-graph context to enforce routing policy — for example, regulatory-sensitive Clinical data routes only to approved/validated models, while low-sensitivity exploratory tasks use lighter-weight, lower-cost models.
• Continuously benchmark and re-tune the model portfolio and routing rules as new models become available, avoiding both over-provisioning expensive frontier models and under-serving tasks that genuinely need them.
AI-Native Data Engineering & Semantic Modeling
• Use AI-assisted and agentic tooling by default across the data lifecycle — pipeline generation, schema/ontology alignment, data-quality scoring, and documentation — not as an occasional accelerator.
• Design conceptual, logical, physical, and semantic models — ontologies, taxonomies, knowledge graphs — that ground both self-service analytics and agentic reasoning, built on and reconciled with the platform's native lineage.
• Build analytics-ready dimensional models and semantic/vector layers that power self-service BI and AI-driven consumption for scientists, statisticians, and business analysts.
Automation & Team Reusability at Scale
• Build reusable agentic accelerators — agent templates, orchestration patterns, evaluation harnesses, model-routing configs — that the broader India team adopts directly to scale AI-native delivery.
• Automate lineage capture, schema-drift detection, and data/agent quality checks so the team spends more time on judgment calls, less on manual upkeep.
• Stand up and maintain the AI/agentic standards platform — the enterprise reference for how agentic solutions, model routing, and knowledge graphs are designed, governed, and delivered across Clinical and Non-Clinical squads.
Governance & Responsible AI
• Implement role-based and attribute-based access control, encryption, and prompt/data-leakage safeguards for agentic systems operating on regulated data.
• Embed data quality, lineage, model/agent risk, and compliance requirements directly into every agentic design and routing decision.
Collaboration & Stakeholder Engagement
• Partner with business SMEs, solution architects, platform/data-catalog teams, and engineering teams to translate Clinical and Non-Clinical needs into scaled agentic solutions.
• Communicate technical trade-offs and AI/agentic architecture recommendations clearly to technical and non-technical stakeholders, and shape the India AI/Agentic roadmap.
Technical Skills & Qualifications Required — AI & Agentic Solutions at Scale (Must-Have, Hands-On)
• 5+ years hands-on building and shipping AI/agentic systems in production: multi-agent orchestration, RAG, tool-calling/function-calling agents, and LLM application frameworks (e.g., LangGraph, LlamaIndex, Semantic Kernel, or equivalent).
• Hands-on experience designing model-routing/orchestration layers (right-model-right-task) across a multi-model portfolio, balancing cost, latency, and accuracy.
• Hands-on experience with vector databases/semantic search, embeddings, and retrieval architectures grounding LLMs in governed data.
• Experience building evaluation frameworks, guardrails, and observability/monitoring for LLM and agentic applications.
Required — Data Platform Integration, Lineage & Semantic Modeling
• Hands-on experience with cloud platforms (AWS) and Unity Catalog-based access control, encryption, and lineage/compliance patterns for regulated data.
• Proven, hands-on semantic data modeling experience: ontologies, taxonomies, knowledge graphs (RDF/OWL or equivalent), built for real Clinical or Non-Clinical data sets.
• Experience integrating agentic/AI systems with enterprise data-catalog and lineage platforms to build knowledge graphs that ground agent reasoning and model-routing decisions, rather than standing up disconnected metadata stores.
• Hands-on experience with cloud-based lakehouse/data platforms (Databricks, or equivalent) as the governed data foundation feeding agentic and AI systems.
Required — AI-Assisted Ways of Working & Governance
• Demonstrated, current practice of using AI-assisted and agentic tools as the default method for analysis, design, and delivery — not occasional use.
• Experience implementing access control, encryption, and responsible-AI/compliance patterns for agentic systems on regulated data.
Preferred
• Certifications or equivalent depth in LLMOps/MLOps platforms, or Databricks Mosaic AI / cloud AI-agent services.
• Exposure to data mesh and data product architecture; API-first design for agent-consumable services.
• Familiarity with GxP / 21 CFR Part 11 compliance and responsible-AI governance in a validated data environment.
Education & Experience Minimum
• Bachelor's degree in Computer Science, Information Systems, or related discipline.
• 12+ years of hands-on data architecture experience, including 5+ years building and shipping AI/agentic solutions to production, with direct experience on Clinical and/or Non-Clinical data.
• Demonstrated track record of AI/agentic architecture work that shipped to production and operates at scale, including integration with an enterprise data platform for lineage and knowledge graphs.
Preferred
• Master's degree or equivalent in a quantitative or engineering discipline.
• Experience in a regulated industry (pharma, life sciences, finance) with GxP or equivalent compliance requirements.
• Prior contribution to clinical or non-clinical data programs supporting IND, NDA, or BLA submissions.
At Lilly, caring is not only what we do for patients. It is how we work. We believe the people who dedicate themselves to making medicines better deserve an environment that makes their lives better too, one where they are supported, respected, and given the space to do their best work. This is not just a policy. It is who we are. Equal Opportunity & Accommodation Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form at https://careers.lilly.com/us/en/workplace-accommodation for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response. Lilly is an EEO/Affirmative Action Employer and does not discriminate on the basis of age, race, colour, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.
#WeAreLilly