Site Reliability Engineer 3

Location IN-KA-Bangalore
ID 2026-10325
Position Type
Full-Time
Employee Type
Regular
Location Type
Remote

The Company

Serving the People Who Serve the People

 

Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive. Granicus has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn.

 

Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 million citizen subscribers power an unmatched Subscriber Network that use our digital solutions to make the world a better place. With comprehensive cloud-based solutions for communications, government website design, meeting and agenda management software, records management, and digital services, Granicus empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, Granicus brings governments closer to the people they serve—driving meaningful change for communities around the globe.


Want to know more? See more of what we do here.

Job Summary

Job Description:

Granicus is seeking a Site Reliability Engineer 3 (SRE) with strong AIOps, automation, and AI proficiency to modernize reliability engineering through observability, intelligent incident response, and responsible AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build scalable, resilient platforms supporting traditional, cloud-native, and AI/ML-powered workloads. The role will also help operationalize AI-enabled SRE practices such as alert intelligence, assisted root-cause analysis, runbook automation, telemetry summarization, and governed self-healing workflows with appropriate human approval and audit controls. You will be expected to implement practical AIOps capabilities across observability, incident response, automation, and operational knowledge workflows, turning AI/ML insights into production-ready reliability improvements.

What Your Impact Will Look Like

  • End-to-end reliability for production systems: on-call, incident response, postmortems, SLO/SLI ownership
  • Building and maintaining observability pipelines (metrics, logs, traces) with AI-assisted anomaly detection layered on top and implementing AIOps pipelines for event ingestion, enrichment, deduplication, correlation, and noise reduction
  • Designing automation that reduces toil — with a clear bias toward AI-augmented runbooks over static scripts and implementing AIOps-driven remediation workflows with approvals, rollback logic, and audit trails
  • Driving alert-noise reduction using correlation/ML techniques, not just threshold tuning
  • Partnering with engineering teams to embed reliability and AI-assisted diagnostics into the SDLC
  • Demonstrated production use of LLMs/AI agents for SRE workflows — e.g., automated log triage, RCA drafting, runbook generation, or incident summarization — not just "I used Copilot to write YAML"
  • Experience building or integrating AIOps capabilities: anomaly detection, alert correlation/clustering, predictive capacity signals including hands-on implementation using observability platforms, ML-based signal processing, incident enrichment, and ChatOps/ticketing integrations
  • Working knowledge of prompt engineering for operational use cases (structured outputs, tool-use/function calling, retrieval-augmented context from runbooks/CMDB)
  • Comfort evaluating AI output critically — can articulate where an LLM's suggested fix or RCA was wrong and why, not just accept it
  • Familiarity with agentic frameworks or MCP-style tool integration (connecting LLMs to ticketing, observability, or ChatOps tools) is a strong plus
  • Understanding of the risk surface of AI-in-production: hallucination in RCA, over-automation risk, human-in-the-loop design for high-blast-radius actions

 

Day-to-day Operations

  • Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration.
  • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups. Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to accelerate diagnosis while validating AI recommendations before action.
  • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting.
  • Design and maintain observability across logs, metrics, and traces, ensuring actionable monitoring and high signal-to-noise alerting.
  • Build and enhance workflows for alerting, anomaly detection, and incident enrichment to reduce noise and improve accuracy. Implement AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment.
  • Develop automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback plans, and auditability. Implement AIOps remediation patterns that connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps.
  • Drive improvements in system reliability, performance, scalability, and resilience through engineering-led initiatives.
  • Partner with engineering teams to improve deployment safety, operational readiness, and production stability.
  • Maintain high-quality runbooks, documentation, and knowledge bases to improve on-call effectiveness and knowledge sharing.
  • Support capacity planning, performance tuning, and SLO-based reliability practices.
  • Apply security, access control, and operational guardrails across systems and automation.
  • Own AIOps implementation from use-case definition through production rollout, including telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.

You Will Love This Job If You Have

· 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments.

· Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP).

· Expert in ELK/OpenSearch, including:

  • Log ingestion (Logstash / Beats)
  • Elasticsearch index design, scaling, and tuning
  • Advanced Kibana querying and debugging
  • Dashboards, alerts, and observability patterns for production systems

· Hands-on experience in logs, metrics, and tracing. Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata.

  • Solid understanding of incident management, RCA, SLOs, and operational best practices.
  • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting. Hands-on implementation experience with AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation.
  • Ability to implement AIOps integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.
  • Experience measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, faster MTTD/MTTR, RCA quality, automation adoption, repeat usage, and business impact linkage.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar.

Preferred Certifications: AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.

 

Job Info

 

  • Shift work – EMEA shifts
  • Remote - INDIA

About Us

Don’t have all the skills/experience mentioned above? At Granicus, we are trying to build diverse, inclusive teams. We do not have degree requirements for most of our roles. If you don’t meet every requirement above but are excited to learn more, we encourage you to apply. We might just be able to find another role that could be a perfect fit!

 

Security and Privacy Requirements

  • Responsible for Granicus information security by appropriately preserving the Confidentiality, Integrity, and Availability (CIA) of Granicus information assets in accordance with the company's information security program.
  • Responsible for ensuring the data privacy of our employees and customers, their data, as well as taking all required privacy training in a timely manner, in accordance with company policies.

 

The Team

  • We are a remote-first company with a globally distributed workforce across the United States, Canada, United Kingdom, India, Armenia, Australia, and New Zealand.

 

The Culture

  • At Granicus, we are building a transparent, inclusive, and safe space for everyone who wants to be
    a part of our journey.
  • A few culture highlights include – Employee Resource Groups to encourage diverse voices
  • Coffee with Mark sessions – Our employees get to interact with our CEO on very important and
    sometimes difficult issues ranging from mental health to work-life balance and current affairs.
  • Microsoft Teams communities focused on wellness, art, furbabies, family, parenting, and more.
  • We bring in special guests from time to time to discuss issues that impact our employee
    population

The Impact

  • We are proud to serve dynamic organizations around the globe that use our digital solutions to make the world a better place — quite literally. We have so many powerful success stories that illustrate how our solutions are impacting the world. See more of our impact here.

Options

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.