DevOps, SRE, Cloud,Applied AIBengaluru, Hyderabad – India
DevOps Specialist Engineer – SRE, Cloud & Applied AI
Location
Bengaluru, Hyderabad – India
Compensation Tier
Mid Level
Category
DevOps, SRE, Cloud,Applied AI
Company Overview
Our client is a leading technology organization focused on building scalable, cloud-native platforms and intelligent software solutions. The organization combines software engineering, cloud technologies, Site Reliability Engineering, and Applied AI to deliver highly resilient and production-ready products at scale.
Position Overview
We are looking for a DevOps Specialist Engineer with strong experience in Site Reliability Engineering, cloud platform engineering, software development, and production operations. The ideal candidate will have hands-on experience operating large-scale distributed systems across Azure, AWS, or GCP, along with exposure to AI/ML, GenAI, LLMOps/MLOps, observability, performance engineering, and cloud cost optimization.
The role requires a strong engineering mindset and the ability to build reliable, secure, scalable, and highly automated production platforms.
Responsibilities
- Design, build, operate, and continuously improve large-scale cloud-native production systems.
- Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
- Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
- Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
- Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
- Develop production-grade observability across metrics, logging, and distributed tracing.
- Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor, and AWS CloudWatch.
- Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
- Support MLOps/LLMOps platforms and AI control-plane capabilities such as model gateways and guardrails.
- Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
- Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
- Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
Skills & Experience
- 6–9 years of experience in DevOps, SRE, Site Reliability Engineering, Platform Engineering, or Software Engineering.
- Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or a related discipline.
- Strong programming experience in one or more of Python, C#/.NET, Go, Java, or Bash.
- Strong hands-on experience with Azure, AWS, or GCP; Azure/AWS preferred.
- Experience with Kubernetes, Docker, Terraform, and cloud-native architectures.
- Strong understanding of CI/CD, GitHub, Azure DevOps (ADO), and ArgoCD.
- Experience with production observability and monitoring using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, CloudWatch, or Azure Monitor.
- Strong understanding of SRE principles, SLIs, SLOs, SLAs, error budgets, incident management, and production on-call operations.
- 3+ years of experience operating or supporting large-scale production systems.
- Experience with AI/ML or GenAI workloads in production and familiarity with Azure OpenAI, AWS Bedrock, or Vertex AI.
- Knowledge of MLOps/LLMOps, MLflow, LangFuse, LangSmith, or equivalent AI/agent orchestration platforms.
- Experience with load/performance testing, capacity planning, autoscaling, and chaos engineering.
Apply for this Role
You must be signed in to apply for this position.
Sign In to Apply