Skip to main content
Technology

Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)

Hyderabad, India

Apply now

Key Responsibilities

- Own platform reliability practices for availability, resilience, latency, and operational efficiency.

- Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.

- Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.

- Lead JFROG Helm chart automation and JFROG images/ACR migration work.

- Support microservices deployment enablement and platform/tooling upgrades.

- Own and optimize monitoring, alerting, observability, and logging stack components:

  Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.

- Support health-check frameworks including Airflow health-check requirements.

- Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.

- Collaborate with architecture and delivery teams on reliability and scalability patterns.

- Lead cloud infrastructure creation, maintenance, governance, and access controls.

- Drive capacity planning, DR planning/exercises, and platform best-practice documentation.

- Support cost management, role enforcement, and license management governance.

- Maintain SOP documentation for established alerts and incident patterns.

Required Qualifications / Must-Have Skills

- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.

- Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.

- Advanced experience with CI/CD engineering and GitHub Actions.

- Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.

- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

- End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

- Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

- Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

- Experience in governance controls: access management, role enforcement, and separation of duties.

- Proven high-severity incident leadership and post-incident reliability improvement execution.

Good-to-Have / Nice-to-Have

- Postgres performance and reliability operations.

- Telecom-scale high-availability systems experience.

Experience Level

Senior to Lead IC (typically 10 to 17 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

What We Offer

- Opportunity to define and scale platform reliability standards.

- High technical ownership and strong cross-functional influence.

- Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Whitefield Road:Intl Tech Park, Navigator Bldg

It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

Job ID R-117416 Date posted 07/30/2026
Apply now

Benefits

Your needs? Met. Your wants? Considered. Take a look at our comprehensive benefits.

  • Paid Time Off
  • Tuition Assistance
  • Medical and dental plans
  • Discounts
  • Training & Development

Learn more about benefits

Our hiring process

Apply Now

Confirm your qualifications align with the job requirements and submit your application.

Assessments

You may be required to complete one or more assessments, depending on the role.

Interview

Get ready to put your best foot forward! More than one interview may be necessary.

Conditional Job Offer

We’ll reach out to discuss a conditional job offer and the next steps to joining the team.

Background Check

Timing is important – complete the necessary actions to proceed with onboarding.

Welcome to the Team!

Congratulations! It’s time to experience #LifeAtATT.

Check your email (and SPAM) throughout the process for important messages and next steps.

Join our talent network

Didn’t find what you were looking for here? Sign up for our job alerts and get the latest AT&T news.

Sign up for the talent network

Don't Miss Out

Join our Talent Network to be the first to know about new job openings, special announcements and behind-the-scenes information.

Skip, I’d rather go straight to the application

AT&T Info and Alerts. Max 12 messages/month Privacy Policy (opens in new window). You may opt-out at anytime by sending STOP to short code 20013. Msg & data rates may apply.

By submitting your information, you acknowledge that you have read our privacy policy (opens in new window) and consent to receive email communication from AT&T for our U.S. Talent Network.