About this role
Meet the Team The Splunk Agent Resilience team is defining the future of AI resilience. Our team provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks. This unified approach empowers our customers to confidently deploy and manage AI-powered applications with enhanced observability and control.
As a Software Engineer on the Data & Scalability Platform team, you will build and operate the systems that store, move, and transform every span, metric, and evaluation result the platform produces. You will own features end to end design, implementation, testing, and production support across our data stores, streaming pipelines, and processing services, and grow into the performance and capacity work that keeps the data plane fast and cost-effective as it scales.
Your Impact
• Design, develop, test, and maintain backend services and data pipelines that operate at high scale. • Build and extend features across the platform's data stores, relational, analytical/columnar, object storage, and caching including schema changes and safe migrations. • Implement and tune the compute and pipelines that move and transform data, including stream processors, writers, and distributed worker fleets. • Solve concrete problems in scalability, reliability, performance, and fault tolerance along the data path. • Improve query and write performance through indexing, query tuning, and hot-path optimization. • Contribute to load testing, profiling, and benchmarking, and turn the results into targeted optimizations. • Add observability instrumentation so the components you build are measurable in production. • Write clean, scalable code and comprehensive tests, primarily in Python, with flexibility to use other backend languages such as Go. • Lead features from technical design through implementation, deployment, and production support. • Debug production issues across data stores, queues, and services, and participate in code reviews, on-call, postmortems, and root-cause analysis. • Drive improvements in engineering practices, automated testing, security, and reliability.
Minimum Qualifications
• Bachelor's degree with 4+ years of related experience, or Master's degree with 2+ years, or PhD with 0 years of relevant software engineering experience. • Strong experience in backend engineering and building production-grade services. • Hands-on experience with large-scale distributed systems, with a focus on scalability, reliability, and performance. • Working experience with at least one database relational or analytical including schema design, query tuning, and safe migrations. • Experience with streaming or queueing systems (e.g., Kafka, RabbitMQ, Pulsar, Kinesis). • Strong proficiency in Python and experience with another backend programming language such as Go, Java, C++, or similar. • Strong fundamentals in system design, distributed systems, APIs, data structures, algorithms, and concurrency. • Experience with modern software engineering practices including CI/CD, automated testing, code reviews, Agile development, and production troubleshooting.
Preferred Qualifications
• Experience with a columnar, analytical, or time-series store (e.g., ClickHouse, Druid, BigQuery, Snowflake). • Experience tuning queue consumers or async task workers for throughput — concurrency, batching, and backpressure. • Exposure to performance work: profiling, benchmarking, or load testing real systems and acting on the measurements. • Experience with cloud-native technologies, containerized environments (Kubernetes), and public cloud platforms (AWS, GCP, or similar). • Experience with observability instrumentation and tooling (OpenTelemetry, metrics, tracing, logs). • Familiarity with multi-tenant systems, quotas, or rate limiting. • Experience with secure coding practices and privacy-by-design. • Strong communication skills and a collaborative approach to design and code review.
Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.
Disclaimer To ensure that we hire the best talent in the right way, we follow a strict hiring process and recently, Cisco has been made aware of fraudulent recruiters claiming to be from the company. Please be advised that any communication from Cisco about careers will:
• be in direct response to an application you have submitted through the company career site • begin with screening or an interview • originate from a Cisco email address, and • be conducted across email, phone, or WebEx Cisco will never make a job offer without conducting an interview process or ask you for money in any way. If you have been requested to apply for a role or have received an offer from a site other than https://careers.cisco.com or cisco.wd5.myworkday.com, do not provide any personal identifying information, including your Aadhaar or other personal identifying number, birth certificate, banking information, driver's license, or passport.
If you are the target of a recruiting scam, consider filing a report with your local law enforcement authorities. Cisco bears no responsibility, and cannot be held liable, for any claims, damages, expenses, or other inconvenience resulting from or in any way connected to recruiting scams.