Open to DevOps · SRE · Cloud TSE roles

Mahendra
Rao.

DevOps & SRE engineer — 11 years turning production chaos into measurable customer trust. Cut MTTR by 35–40% across two enterprise SaaS platforms. Currently building for the observability + AI-ops era.

scroll to explore

// 01 — About

Eleven years
of production trust.

The through-line: I've spent 11 years standing between production systems and the customers who depend on them. Every role — telecom IVR, enterprise APIs, Fortune 500 SaaS, OTA vehicle platforms, cloud observability — has been the same underlying job: find the truth in noisy systems, own it end-to-end, and turn what's broken into what's measurable.

What that looks like in numbers: reduced customer MTTR by 35% at Middleware.io across 50+ enterprise observability accounts, cut MTTR another 40% at Ola Electric while improving deployment velocity 40% and eliminating 80% of manual release steps, and scaled a 24/7 L3 SRE team from 7 to 20 engineers. At Sprinklr, I ran technical accounts for Apple, Microsoft, Dell, Samsung, and UBS at 95%+ CSAT — including an SSO/SAML migration for 100,000+ enterprise users.

Where I'm heading: deeper into SRE, platform engineering, and cloud solutions — with a particular interest in how observability, incident response, and identity are reshaping around AI-driven workloads. My open-source project ClawBridge is where I'm exploring that intersection publicly. Currently in Bengaluru, open to remote and on-site roles.

11+
Years in production
35–40%
MTTR reduction across roles
100K+
Users migrated to SSO/SAML
7 → 20
L3 SRE team scaled

Where I've delivered

Enterprise SaaS B2B APIs Cloud Observability Identity / SSO OTA / IoT platforms Fortune 500 accounts

How I operate

End-to-end ownership Root cause first Runbook everything SLO / SLI driven Product feedback loops Mentor the next line

What comes next

Senior SRE / Platform Cloud TSE / Solutions Observability for AI Incident-response tooling Open-source contribution

// 02 — Tech Stack

What I
build with.

Infrastructure
Kubernetes Docker GCP (GKE · Cloud Run) Azure Terraform Helm Linux
CI/CD & Delivery
GitHub Actions Azure DevOps Jenkins GitOps Infrastructure as Code
Observability
OpenTelemetry Prometheus Grafana Middleware.io Dynatrace New Relic Splunk ELK
Languages
Python Bash TypeScript Node.js YAML JSON
// production cloud: GCP & Azure · currently learning: AWS (SAA-C03)

// 03 — Experience

Where I've
built my craft.

March 2025 — December 2025
Middleware.io
Technical Support Engineer · Cloud Observability

Owned end-to-end observability support for 50+ global enterprise SaaS customers on a modern OpenTelemetry-native platform — the direct engineering-and-customer interface for a fast-growing observability product.

The situation
Customers running distributed workloads across Node.js, Python, Kubernetes, and Cloud Run were seeing intermittent latency, memory pressure, and gaps in their traces. Ticket volume was outpacing engineering's ability to root-cause quickly.
What I did
Reproduced customer environments locally, instrumented workloads with OpenTelemetry, and used Prometheus + Grafana to isolate root causes. Wrote runbooks for the recurring failure modes. Fed patterns back to engineering as roadmap-shaping product feedback.
What changed
MTTR down 35% across the account portfolio. Repeat-issue volume dropped as runbooks became customer self-service. Two of my product-feedback threads shipped as features in OpsAI and Continuous Profiling.
What's next
Deeper into the AI-observability space — instrumentation for LLM applications, span-drop diagnostics, and support tooling for the next generation of workloads. Building OTel Debug Companion as a personal follow-through.
OpenTelemetry APM Distributed Tracing Prometheus Grafana Kubernetes GCP Cloud Run Node.js Python
April 2023 — February 2025
Independent Practice & Full-Time Parent
Deliberate career break · Hands-on upskilling in modern DevOps

Took a planned break to be present for family — and used the time deliberately to move my stack forward from the older telco/enterprise-integration era into cloud-native, container-orchestrated, GitOps-driven engineering.

The situation
My prior decade was strong on enterprise APIs, incident management, and customer engineering — but light on the modern container / IaC / cloud-CI/CD tooling that senior platform roles now expect. Rather than fake it, I set out to genuinely earn it.
What I did
Built end-to-end pipelines with GitHub Actions deploying to GCP Cloud Run. Ran Prometheus/Grafana with Slack-integrated alerting. Practiced Terraform, Docker Compose, and Kubernetes locally. Earned Power BI Data Analyst; started CKA / CKAD / GCP DevOps Engineer / AWS SAA-C03.
What changed
Walked into Middleware.io on the strength of that self-directed practice — and immediately delivered the 35% MTTR improvement. The upskilling wasn't cosmetic; it was the reason the next role happened.
What's next
Finishing the cert stack (AWS SAA-C03 in progress) and moving hands-on time toward what I'm actually building — ClawBridge, OTel Debug Companion, and public writing about incident response and observability craft.
GitHub Actions GCP Cloud Run Kubernetes Docker Terraform Prometheus / Grafana CKA · CKAD (in progress)
April 2022 — April 2023
Ola Electric
L3 Product Support Engineer · SRE / DevOps

Owned production reliability for the OTA vehicle software platform at a hypergrowth EV company — the pipeline that pushed firmware and app updates to a live fleet of electric two-wheelers.

The situation
A young platform, a growing fleet, and a small 24/7 L3 support team (7 engineers). Deployments were largely manual, MTTR was inconsistent, and the on-call load was crushing the team as customer count kept climbing.
What I did
Rebuilt the CI/CD pipeline on Azure DevOps with proper stage gates. Stood up observability across Grafana, Dynatrace, and Splunk. Standardized RCA and post-mortem process. Formalized on-call rotations, SLAs, training material, and a knowledge base. Managed API Gateway integrations via Apigee.
What changed
Deployment velocity +40%. Manual release steps −80%. MTTR down 40% through automation and structured post-mortems. Team scaled from 7 → 20 engineers with an on-call culture that survived my exit.
What's next
The lessons — on-call sustainability, RCA discipline, automation-as-culture — carry directly into how I think about senior SRE and platform work today. This role is why I take runbook craft as seriously as I do.
Azure DevOps CI/CD Grafana Dynatrace Splunk Apigee SRE Incident Management Team scaling
May 2019 — April 2022
Sprinklr
Technical Account Manager · Enterprise Identity & Integrations

Technical SPOC for a Fortune 500 enterprise portfolio — Apple, Microsoft, Dell, Samsung, UBS — on the Sprinklr CXM platform. The role sat right where API integration depth met executive-level customer trust.

The situation
Enterprise customers were consolidating identity providers, migrating off legacy community platforms (GetSatisfaction), and demanding zero-downtime B2B API integrations across Salesforce, Slack, Zendesk, and Adobe Analytics. Any misstep would land in a Fortune 500 exec inbox.
What I did
Architected the GetSatisfaction → Sprinklr Communities migration with SSO/SAML for 100,000+ enterprise users. Ran the B2B integration work end-to-end (OAuth2/REST contract validation, staged cutovers, rollback plans). Led critical-incident RCAs and mentored junior engineers on customer-facing craft.
What changed
Sustained 95%+ CSAT across the Fortune 500 portfolio. The community migration was the largest of its kind on the platform that year — zero downtime, zero identity failures. Multiple accounts renewed multi-year on the strength of that delivery.
What's next
The identity/auth muscle from here (OAuth2, OIDC, SAML, SSO, API Gateway) is exactly what powers my current interest in Okta / Auth0-style identity roles and in supporting enterprise SaaS at that scale again.
Salesforce API OAuth2 / OIDC SAML / SSO B2B Integrations Enterprise SaaS Fortune 500 95%+ CSAT
2014 — 2019
Plivo · Movate (CSS Corp) · IBM · Aditya Birla Group
Enterprise APIs · IVR / VoIP · Production support

Five years of foundational work across telecom (Plivo), IT services (Movate, IBM), and enterprise conglomerates (Aditya Birla) — supporting APIs, IVR / VoIP platforms, and production incident workflows at scale.

What it built
The systematic debugging instinct — subdivide a problem, follow the packet, don't guess — and the customer-facing composure that later made Fortune 500 accounts possible. This is the layer everything since has been built on.
What it means now
Whenever I sit down to debug a modern distributed system, I still use the same first-principles approach I learned here: find the source of truth, verify at each layer, and never trust a symptom. Tools change; the discipline doesn't.
Enterprise APIs IVR / VoIP Production Troubleshooting Telecom Foundational years

// 04 — Skills

Tools of
the trade.

📡
Observability & APM
End-to-end telemetry — metrics, logs, traces, RUM. Deep expertise in OpenTelemetry, Prometheus, Grafana, Dynatrace, New Relic, Splunk, and ELK at enterprise scale.
☁️
Google Cloud Platform
GCE, GKE, Cloud Run, Cloud Logging & Monitoring, IAM. Comfortable across the full Compute stack and cloud-native architecture patterns.
🐳
Kubernetes
Pod debugging, resource constraints, HPA, PVCs, node affinity, DaemonSets. Skilled at diagnosing cluster-level failures and translating them into stakeholder-legible narratives.
⚙️
CI/CD & DevOps
GitHub Actions, Azure DevOps, Jenkins. Designed pipelines that cut deployment time by 40% and eliminated 80% of manual release steps at production scale.
🐧
Linux Systems
Performance analysis, process management, filesystem diagnostics, systemd, cgroups, kernel parameters. Comfortable operating in production at any hour.
🌐
Networking
TCP/IP, DNS, HTTP/S, TLS, VPCs, load balancers, firewalls. Can trace a packet from browser to backend — and explain exactly where and why it got dropped.
🐍
Python & Scripting
Python for automation, API testing, monitoring integrations, and tooling. Bash for ops scripts, diagnostic pipelines, and environment bootstrapping.
🔐
Security & IAM
OAuth2, OIDC, SAML, SSO. Experience with API Gateway configurations (Apigee), enterprise SSO migrations for 100K+ users, and zero-downtime auth integrations.
🤝
Customer Success & SRE
QBRs, adoption metrics, churn prevention, escalation leadership. Scaled L3 SRE teams from 7→20 engineers. Proven track record of turning at-risk accounts into long-term champions.

// 05 — Projects & Case Studies

Where the work
lives publicly.

Flagship · Open Source · Live MVP github.com/MahendraRao/clawbridge

ClawBridge — a friendly bridge to OpenClaw.

My most current, publicly-visible engineering work — and my genuine entry point into the AI-tooling and developer-experience space. Built end-to-end with a modern JavaScript/TypeScript stack, designed to lower the technical friction of using OpenClaw for non-developer users.

The problem
OpenClaw is powerful but has a steep setup curve — install dependencies, configure providers (OpenAI / Anthropic / Ollama / custom), validate the environment, launch. Every one of those steps is a place a non-developer gives up.
The approach
React + Vite + TypeScript frontend paired with a local Express + Node.js API. Handles install, system diagnostics, provider orchestration, and first-run onboarding — all through a UI a non-coder can navigate.
Current state
Working MVP today. Real installation flow, working provider integration (OpenAI + Anthropic + Ollama), local diagnostics. Public on GitHub — architecture and code visible for review.
What's next
Electron packaging for cross-platform desktop, OpenTelemetry instrumentation for observability, and Docker / Kubernetes deployment paths. Actively developed — happy to walk anyone through the code.
GitHub · HTML / Three.js Live
This Portfolio

A production-grade personal portfolio site with a Three.js 3D hero, scroll-triggered animations, amber/charcoal dark theme, and full mobile responsiveness. Built from scratch — no templates.

Three.jsCSS AnimationsVercel
↗ github.com/MahendraRao/mahendra-portfolio
GitHub · Three.js Live
3D Humpty — Interactive Character

An interactive 3D character experience built with Three.js. A dancing Humpty Dumpty you can control — built to sharpen my 3D rendering and real-time animation skills in the browser.

Three.js3D AnimationInteractive
↗ github.com/MahendraRao/3d-humpty
GitHub · HTML / JS ★ 5 Stars
Consciousness

An exploratory creative coding experiment — pushing the boundaries of what pure HTML and JavaScript can express. One of my most starred personal projects.

Creative CodingHTML CanvasGenerative
↗ github.com/MahendraRao/consciousness
Coming Soon · Python / OTel In Progress
OTel Debug Companion

A CLI tool that interrogates your OpenTelemetry pipeline — checking collector health, trace ingestion rates, span drop rates, and SDK configuration. The runbook I wished existed during enterprise escalations at Middleware.

OpenTelemetryPythonObservabilityCLI
↗ Coming soon on GitHub
Coming Soon · GitHub Actions In Progress
GCP Cloud Run Pipeline Template

A production-ready GitHub Actions workflow template for deploying to GCP Cloud Run — with secret management, smoke tests, rollback on failure, and Slack notifications. Built during upskilling; polishing for open source.

GitHub ActionsGCPCloud RunDocker
↗ Coming soon on GitHub
Case Study · Middleware.io
Zero-to-OpenTelemetry Enterprise Onboarding

Designed a structured onboarding for enterprise customers adopting OpenTelemetry from scratch — instrumentation guides, debug runbooks, and troubleshooting playbooks for Go, Python, and Node.js. Cut time-to-value from weeks to days.

OpenTelemetryOnboardingTime-to-Value
Case Study · Ola Electric
Escalation Framework Built from Ground Zero

Built a tiered escalation framework and knowledge infrastructure during hypergrowth. Defined SLAs, trained L1/L2 agents, and created engineering feedback loops that measurably reduced repeat incidents across the product lifecycle.

Escalation DesignSLATraining
Case Study · Sprinklr
Fortune 500 Account Success Program

Ran strategic success for Fortune 500 accounts on the Sprinklr CXM platform. Delivered impactful QBRs, drove feature adoption, and acted as the technical-business bridge during critical escalations — converting at-risk accounts into multi-year renewals.

QBRAdoptionRenewal

// 06 — Writing

Thoughts on
the craft.

Coming Soon
OpenTelemetry in Production: What Nobody Tells You

Lessons from instrumenting real production systems — the gotchas, the sampling edge cases, and why your traces are probably lying to you.

Read more →
Coming Soon
The CI/CD Pipeline That Saved 80% of Our Manual Work

How we redesigned the OTA software deployment pipeline at Ola Electric — and what most teams get wrong about release velocity vs. release safety.

Read more →
Coming Soon
The Kubernetes Debugging Playbook I Wish Existed on Day One

A field-tested guide to pod failures, resource starvation, and network black holes — written from a support engineer's perspective, not a textbook author's.

Read more →

// 07 — Contact

Let's build
something together.

Open to DevOps Engineer, SRE, Cloud TSE, Solutions Engineer, and Platform Engineering roles — especially in observability, cloud infrastructure, or enterprise SaaS. Based in Bengaluru · Open to remote & hybrid.

nkneelkumar [at] gmail [dot] com
↗ LinkedIn ↗ GitHub ↗ Twitter / X ↗ Resume (PDF)