DevOps & SRE engineer — 11 years turning production chaos into measurable customer trust. Cut MTTR by 35–40% across two enterprise SaaS platforms. Currently building for the observability + AI-ops era.
// 01 — About
The through-line: I've spent 11 years standing between production systems and the customers who depend on them. Every role — telecom IVR, enterprise APIs, Fortune 500 SaaS, OTA vehicle platforms, cloud observability — has been the same underlying job: find the truth in noisy systems, own it end-to-end, and turn what's broken into what's measurable.
What that looks like in numbers: reduced customer MTTR by 35% at Middleware.io across 50+ enterprise observability accounts, cut MTTR another 40% at Ola Electric while improving deployment velocity 40% and eliminating 80% of manual release steps, and scaled a 24/7 L3 SRE team from 7 to 20 engineers. At Sprinklr, I ran technical accounts for Apple, Microsoft, Dell, Samsung, and UBS at 95%+ CSAT — including an SSO/SAML migration for 100,000+ enterprise users.
Where I'm heading: deeper into SRE, platform engineering, and cloud solutions — with a particular interest in how observability, incident response, and identity are reshaping around AI-driven workloads. My open-source project ClawBridge is where I'm exploring that intersection publicly. Currently in Bengaluru, open to remote and on-site roles.
// 02 — Tech Stack
// 03 — Experience
Owned end-to-end observability support for 50+ global enterprise SaaS customers on a modern OpenTelemetry-native platform — the direct engineering-and-customer interface for a fast-growing observability product.
Took a planned break to be present for family — and used the time deliberately to move my stack forward from the older telco/enterprise-integration era into cloud-native, container-orchestrated, GitOps-driven engineering.
Owned production reliability for the OTA vehicle software platform at a hypergrowth EV company — the pipeline that pushed firmware and app updates to a live fleet of electric two-wheelers.
Technical SPOC for a Fortune 500 enterprise portfolio — Apple, Microsoft, Dell, Samsung, UBS — on the Sprinklr CXM platform. The role sat right where API integration depth met executive-level customer trust.
Five years of foundational work across telecom (Plivo), IT services (Movate, IBM), and enterprise conglomerates (Aditya Birla) — supporting APIs, IVR / VoIP platforms, and production incident workflows at scale.
// 04 — Skills
// 05 — Projects & Case Studies
My most current, publicly-visible engineering work — and my genuine entry point into the AI-tooling and developer-experience space. Built end-to-end with a modern JavaScript/TypeScript stack, designed to lower the technical friction of using OpenClaw for non-developer users.
A production-grade personal portfolio site with a Three.js 3D hero, scroll-triggered animations, amber/charcoal dark theme, and full mobile responsiveness. Built from scratch — no templates.
↗ github.com/MahendraRao/mahendra-portfolioAn interactive 3D character experience built with Three.js. A dancing Humpty Dumpty you can control — built to sharpen my 3D rendering and real-time animation skills in the browser.
↗ github.com/MahendraRao/3d-humptyAn exploratory creative coding experiment — pushing the boundaries of what pure HTML and JavaScript can express. One of my most starred personal projects.
↗ github.com/MahendraRao/consciousnessA CLI tool that interrogates your OpenTelemetry pipeline — checking collector health, trace ingestion rates, span drop rates, and SDK configuration. The runbook I wished existed during enterprise escalations at Middleware.
↗ Coming soon on GitHubA production-ready GitHub Actions workflow template for deploying to GCP Cloud Run — with secret management, smoke tests, rollback on failure, and Slack notifications. Built during upskilling; polishing for open source.
↗ Coming soon on GitHubDesigned a structured onboarding for enterprise customers adopting OpenTelemetry from scratch — instrumentation guides, debug runbooks, and troubleshooting playbooks for Go, Python, and Node.js. Cut time-to-value from weeks to days.
Built a tiered escalation framework and knowledge infrastructure during hypergrowth. Defined SLAs, trained L1/L2 agents, and created engineering feedback loops that measurably reduced repeat incidents across the product lifecycle.
Ran strategic success for Fortune 500 accounts on the Sprinklr CXM platform. Delivered impactful QBRs, drove feature adoption, and acted as the technical-business bridge during critical escalations — converting at-risk accounts into multi-year renewals.
// 06 — Writing
Lessons from instrumenting real production systems — the gotchas, the sampling edge cases, and why your traces are probably lying to you.
Read more →How we redesigned the OTA software deployment pipeline at Ola Electric — and what most teams get wrong about release velocity vs. release safety.
Read more →A field-tested guide to pod failures, resource starvation, and network black holes — written from a support engineer's perspective, not a textbook author's.
Read more →// 07 — Contact
Open to DevOps Engineer, SRE, Cloud TSE, Solutions Engineer, and Platform Engineering roles — especially in observability, cloud infrastructure, or enterprise SaaS. Based in Bengaluru · Open to remote & hybrid.
nkneelkumar [at] gmail [dot] com