AI Agentic Ops SRE

Tags: ,

Role Title: AI Agentic Ops SRE

Employer: Leading Financial Technology Firm

Required Experience: 6+ Years

Location: Pan India

Date published: 17 July 2026

A Leading Financial Technology Firm is seeking a highly technical AI Agentic Ops SRE to deploy and operate autonomous AI architectures across its production networks. In this cutting-edge platform reliability role, you will take absolute responsibility for running AI systems that think, plan, and execute task strings with minimal human intervention. Furthermore, you will build automated self-healing frameworks to safeguard cloud ecosystems from service disruption. Consequently, this position is vital for driving operational excellence and scaling AI-driven site reliability engineering.

The AI Agentic Ops SRE must combine deep site reliability infrastructure expertise with an advanced understanding of multi-agent monitoring tools. Working within a high-volume data ecosystem, you will utilize Large Language Models to diagnose code defects, trace network latencies, and automate chaos engineering workflows. Therefore, the group is looking for an automation-focused engineer who resolves Tier-3 infrastructure incidents comfortably under tight operational timelines. If you want to engineer tomorrow’s autonomous cloud operations, this position is for you.

Key Responsibilities

  • Deploy, optimize, and manage autonomous AI agents explicitly configured for cloud computing platforms and production operations.
  • Provide expert Tier-3 production support for critical cloud infrastructure systems and automated AI agent data pipelines.
  • Participate actively in on-call engineering rotations, troubleshoot production incidents, and perform deep root cause analysis (RCA).
  • Leverage LLM-based agents to parse server logs, correlate application traces, and automatically generate incident post-mortems.
  • Deploy continuous chaos-testing agents across networks to proactively identify infrastructure and AI system vulnerabilities.
  • Monitor autonomous AI agent performance metrics, platform reliability indices, token utilization rates, and overall system health.
  • Integrate specialized monitoring systems using Galileo AI and LangSmith for prompt tracing, cost tracking, and latency checks.
  • Utilize Galileo and LangSmith platforms to lead RAG evaluations, hallucination detection tests, and agent performance analysis.
  • Configure and maintain advanced monitoring dashboards and distributed tracing networks via Loki, Grafana, Tempo, and Mimir.
  • Collaborate with platform squads to establish automated alerting boundaries, metric metrics, and continuous system optimizations.

Requirements and Qualifications

  • Bachelor’s degree in Computer Science, Artificial Intelligence, Software Engineering, Data Science, or an allied technical stream.
  • 6+ Years of successful experience operating as an SRE, DevOps Architect, or Cloud Systems Automation Specialist.
  • Proven core expertise with observability tools for multi-agent workflows, handling LangChain, LangGraph, and LangSmith.
  • Hands-on production experience navigating telemetry stacks including Loki, Grafana, Tempo, and Mimir/Prometheus.
  • Deep structural understanding of cloud infrastructure components, RAG evaluation methods, and hallucination management rules.
Apply Now
Note: Submitting your resume does not mean it will be directly shared with the employer. Your consent will be sought before any information is shared.