You are viewing a preview of this job. Log in or register to view more details about this job.

AI-Enabled Platform SRE Engineer

AI-Enabled Platform/SRE Engineer

 

Location: Dallas, TX / Scottsdale, AZ (Hybrid)

Assessment Required: ROPES Assessment (Mandatory)

Interview Process: Onsite Interview in Dallas, TX or Scottsdale, AZ

 

Position Overview

We are seeking a Senior Kubernetes-focused Site Reliability Engineer (SRE) with strong expertise in cloud automation, Kubernetes platform engineering, and software development. The ideal candidate will leverage AI/LLMs to automate operations, improve platform reliability, and build scalable cloud-native infrastructure.

Candidates must have strong hands-on experience with Kubernetes, GCP, Python, Java, and Site Reliability Engineering (SRE).

Required Skills

5+ years of hands-on experience in Site Reliability Engineering (SRE)
5+ years of experience with Kubernetes Platform Engineering
• Strong experience with Google Kubernetes Engine (GKE) and Rancher RKE2
• Strong experience with Google Cloud Platform (GCP)
• Experience with Terraform, Helm, GitHub, and CI/CD
5+ years of programming experience using Python and Java
• Experience with Node.js for automation and integrations (Preferred)
• Strong experience with REST APIs, GraphQL, Apigee, and Apigee X
• Experience with Splunk, Grafana, Datadog, and AppDynamics
• Experience implementing AI-driven Operations (AIOps) using Gemini, Llama, Mistral, Qwen, or similar LLMs
• Strong troubleshooting, automation, and production support experience

 

Key Responsibilities

• Build automation and operational tools using Python, Java, and Node.js
• Develop AI-powered operational workflows using LLMs for alert analysis, incident response, and automation
• Design and maintain Kubernetes environments across GKE and Rancher RKE2
• Implement cloud automation using Terraform, Helm, and CI/CD pipelines
• Improve platform reliability, scalability, and operational efficiency
• Design and support API reliability using Apigee, REST APIs, and GraphQL
• Implement traffic routing, canary deployments, failover strategies, and disaster recovery solutions
• Build monitoring, logging, dashboards, and alerting using Splunk, Grafana, Datadog, and AppDynamics
• Support multi-cluster Kubernetes environments and active-active deployments
• Collaborate with cross-functional teams to improve platform stability and operational excellence

 

Preferred Qualifications

• Experience with enterprise Kubernetes platforms
• Experience building AI-powered operational automation
• Experience with distributed systems and cloud-native architectures
• Strong scripting and automation skills
• Excellent communication and collaboration skills

 

Recruitment Notice

This position is with our client, Randstad, and we (ThinqSpot Inc.) are the authorized third-party recruiting partner for this opportunity. Applicants must apply through ThinqSpot Inc. We will coordinate interviews and guide candidates throughout the hiring process.

 

About ThinqSpot Inc.

ThinqSpot Inc. is a staffing agency specializing in connecting top technology professionals with our clients. In addition to supporting our clients' hiring needs, we also recruit for our internal projects and direct clients.

We provide highly qualified technology professionals to meet our clients' business requirements and do not share any candidate's information with any third party without the candidate's informed written consent.

Important: ThinqSpot Inc. never charges candidates any upfront fees or recruitment costs at any stage of the hiring process.