☁️ DevOps Engineer — Site reliability engineer

📍 Location: Noida | 🕐 Full-Time | 🧭 Experience: 1–3 years

About Testmu AI🚀
LambdaTest is a high-growth SaaS platform powering millions of test executions globally with 100% YoY traffic growth. Join our Platform Engineering team to own and scale the cloud infrastructure that keeps our systems fast, reliable, and production-ready.

The Role: What You'll Do
  • This is a genuine DevOps/SRE role — not config management or ticket ops. You'll write code, own incidents end-to-end, and build custom automation solutions on live production systems.
  • Cloud Operations: Own and scale services on AWS — EKS, EC2, SQS, ECR, Route 53, ALB/NLB.
  • Kubernetes & Containers: Manage production workloads on EKS; implement Karpenter and KEDA for event-driven and node autoscaling.
  • CI/CD: Build and optimize pipelines using Jenkins, Docker Buildx, ECR caching, ArgoCD, and Helm.
  • Observability & Incident Management: Own APM, monitoring, and production debugging using New Relic, Sumo Logic, Prometheus, or Grafana.
  • Automation & Scripting: Write custom code for logical and automation solutions — not just run commands on machines.
  • IaC: Use Terraform for provisioning; understand state management and concurrent apply risks in team environments.

You'll Thrive Here If You Have | Must-Haves:
  • Experience: 1–3 years of hands-on DevOps or Cloud Infrastructure experience.
  • Cloud: Production experience on AWS — EKS, SQS, ECR, Route 53, ALB/NLB.
  • Containers: Docker and Kubernetes in production environments.
  • Observability: Hands-on with at least one APM tool — New Relic, Sumo Logic, Prometheus, or Grafana.
  • Incident Management: Real production incident experience — must articulate root cause, not just symptoms.
  • Coding: Genuine programming ability in any language — this role builds custom solutions, not just runs CLI commands.
  • Mindset: Bridges dev and DevOps thinking; can explain why architectural choices were made, not just what was done.
Good to Have:
  • Golang or Java — backend coding ability is a strong plus.
  • KEDA and Karpenter — event-driven and node autoscaling experience.
  • ArgoCD, Helm, or Istio — GitOps and service mesh exposure.
  • Kafka or SQS — messaging and queue configuration knowledge.
  • System design awareness at the services layer — how services interact, not just surface-level ops.

What We Offer
Direct exposure to production systems at scale, ownership from day one, and a clear growth path in Platform and SRE Engineering.


Apply now to build infrastructure that scales.

Required Skills

AZURE Kafka Redis Golang