Tags: Galileo AI, LangSmith
Role Title: AI Agentic Ops SRE
Employer: Leading Financial Technology Firm
Required Experience: 6+ Years
Location: Pan India
Date published: 17 July 2026
A Leading Financial Technology Firm is seeking a highly technical AI Agentic Ops SRE to deploy and operate autonomous AI architectures across its production networks. In this cutting-edge platform reliability role, you will take absolute responsibility for running AI systems that think, plan, and execute task strings with minimal human intervention. Furthermore, you will build automated self-healing frameworks to safeguard cloud ecosystems from service disruption. Consequently, this position is vital for driving operational excellence and scaling AI-driven site reliability engineering.
The AI Agentic Ops SRE must combine deep site reliability infrastructure expertise with an advanced understanding of multi-agent monitoring tools. Working within a high-volume data ecosystem, you will utilize Large Language Models to diagnose code defects, trace network latencies, and automate chaos engineering workflows. Therefore, the group is looking for an automation-focused engineer who resolves Tier-3 infrastructure incidents comfortably under tight operational timelines. If you want to engineer tomorrow’s autonomous cloud operations, this position is for you.
Key Responsibilities
- Deploy, optimize, and manage autonomous AI agents explicitly configured for cloud computing platforms and production operations.
- Provide expert Tier-3 production support for critical cloud infrastructure systems and automated AI agent data pipelines.
- Participate actively in on-call engineering rotations, troubleshoot production incidents, and perform deep root cause analysis (RCA).
- Leverage LLM-based agents to parse server logs, correlate application traces, and automatically generate incident post-mortems.
- Deploy continuous chaos-testing agents across networks to proactively identify infrastructure and AI system vulnerabilities.
- Monitor autonomous AI agent performance metrics, platform reliability indices, token utilization rates, and overall system health.
- Integrate specialized monitoring systems using Galileo AI and LangSmith for prompt tracing, cost tracking, and latency checks.
- Utilize Galileo and LangSmith platforms to lead RAG evaluations, hallucination detection tests, and agent performance analysis.
- Configure and maintain advanced monitoring dashboards and distributed tracing networks via Loki, Grafana, Tempo, and Mimir.
- Collaborate with platform squads to establish automated alerting boundaries, metric metrics, and continuous system optimizations.
Requirements and Qualifications
- Bachelor’s degree in Computer Science, Artificial Intelligence, Software Engineering, Data Science, or an allied technical stream.
- 6+ Years of successful experience operating as an SRE, DevOps Architect, or Cloud Systems Automation Specialist.
- Proven core expertise with observability tools for multi-agent workflows, handling LangChain, LangGraph, and LangSmith.
- Hands-on production experience navigating telemetry stacks including Loki, Grafana, Tempo, and Mimir/Prometheus.
- Deep structural understanding of cloud infrastructure components, RAG evaluation methods, and hallucination management rules.