AI Agentic Ops SRE

Tags: ,

Role Title: AI Agentic Ops SRE

Employer: Leading Financial Technology Firm

Required Experience: 6+ Years

Location: Pan India

Date published: 17 July 2026

A Leading Financial Technology Firm is seeking a highly technical AI Agentic Ops SRE to deploy and operate autonomous AI architectures across its production networks. In this cutting-edge platform reliability role, you will take absolute responsibility for running AI systems that think, plan, and execute task strings with minimal human intervention. Furthermore, you will build automated self-healing frameworks to safeguard cloud ecosystems from service disruption. Consequently, this position is vital for driving operational excellence and scaling AI-driven site reliability engineering.

The AI Agentic Ops SRE must combine deep site reliability infrastructure expertise with an advanced understanding of multi-agent monitoring tools. Working within a high-volume data ecosystem, you will utilize Large Language Models to diagnose code defects, trace network latencies, and automate chaos engineering workflows. Therefore, the group is looking for an automation-focused engineer who resolves Tier-3 infrastructure incidents comfortably under tight operational timelines. If you want to engineer tomorrow’s autonomous cloud operations, this position is for you.

Key Responsibilities

  • Deploy, optimize, and manage autonomous AI agents explicitly configured for cloud computing platforms and production operations.
  • Provide expert Tier-3 production support for critical cloud infrastructure systems and automated AI agent data pipelines.
  • Participate actively in on-call engineering rotations, troubleshoot production incidents, and perform deep root cause analysis (RCA).
  • Leverage LLM-based agents to parse server logs, correlate application traces, and automatically generate incident post-mortems.
  • Deploy continuous chaos-testing agents across networks to proactively identify infrastructure and AI system vulnerabilities.
  • Monitor autonomous AI agent performance metrics, platform reliability indices, token utilization rates, and overall system health.
  • Integrate specialized monitoring systems using Galileo AI and LangSmith for prompt tracing, cost tracking, and latency checks.
  • Utilize Galileo and LangSmith platforms to lead RAG evaluations, hallucination detection tests, and agent performance analysis.
  • Configure and maintain advanced monitoring dashboards and distributed tracing networks via Loki, Grafana, Tempo, and Mimir.
  • Collaborate with platform squads to establish automated alerting boundaries, metric metrics, and continuous system optimizations.

Requirements and Qualifications

  • Bachelor’s degree in Computer Science, Artificial Intelligence, Software Engineering, Data Science, or an allied technical stream.
  • 6+ Years of successful experience operating as an SRE, DevOps Architect, or Cloud Systems Automation Specialist.
  • Proven core expertise with observability tools for multi-agent workflows, handling LangChain, LangGraph, and LangSmith.
  • Hands-on production experience navigating telemetry stacks including Loki, Grafana, Tempo, and Mimir/Prometheus.
  • Deep structural understanding of cloud infrastructure components, RAG evaluation methods, and hallucination management rules.
Apply Now
Note: Submitting your resume does not mean it will be directly shared with the employer. Your consent will be sought before any information is shared.
tilted-delta.com
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.