Site Reliability Engineering | API Platforms

Keeping enterprise API platforms reliable, observable, and ready for production.

I support large API and integration platforms with a focus on service reliability, incident response, monitoring, automation, release readiness, performance improvement, and disaster recovery planning.

API Estate 100+ Services
Monitor
Respond
Automate
Recover
SRE Focus Reliability, readiness, and uptime
6+ Years in SRE, production support, and platform operations
100+ API services supported through reliability and resiliency practices
3x REST notification throughput improvement from 500/min to 1500/min
RTB Run-the-bank ownership across incidents, monitoring, releases, and controls

A site reliability engineer focused on stable services and practical operations.

Hema Kamineni has experience supporting enterprise API platforms at Barclays and coordinating work across BTB, RTB, risk, cyber security, infrastructure, service managers, vendors, and partner teams. Her work is centered on keeping services healthy, improving observability, handling incidents, supporting releases, reducing repeat issues, and making sure resiliency and disaster recovery plans are ready when needed.

SRE areas Hema works across

A practical mix of production support, service health, monitoring, incident response, automation, platform ownership, and resiliency management.

Service Reliability

Supports API platforms with a focus on uptime, stability, service health, and reducing repeat incidents.

Observability & Monitoring

Uses tools like AppDynamics, Kibana, Splunk, ThousandEyes, and GenEos to track issues and improve visibility.

Incident Response

Works through P1/P2 incidents, major incident follow-ups, root-cause analysis, and preventive actions.

Automation & Runbooks

Improves support handoffs with runbooks, KT material, restart procedures, and Jenkins-driven operations.

Release Readiness

Coordinates upcoming releases and changes so deployments are planned, reviewed, and less risky.

Disaster Recovery & Resiliency

Supports PARP, DIRP, DCR, and SRP work as part of a broader reliability and service continuity program.

Professional experience

Jan 2021 - Present

Senior Application Support Administrator, API Platform RTB / SRE

Barclays | Whippany, NJ

Supports reliability for an enterprise API estate covering 100+ services. Handles service health, monitoring, incident follow-ups, operational readiness, release coordination, resiliency planning, QA finding closure, and cyber/NPA remediation work across multiple teams.

Nov 2018 - Dec 2020

Production Support Analyst

Barclays via AppLab Systems Inc. | Wilmington, DE

Supported Java and API-based applications for Credit Cards and Business Banking. Worked with SOA, SOAP, WSDL, REST, CI/CD deployments, monitoring, resiliency activities, RCA documentation, and performance fixes.

Jan 2018 - Oct 2018

Application Support Analyst

Dizer Corp | Painesville, OH

Supported distributed applications using Unix, SQL, Java, and Tomcat. Performed daily health checks, led incident calls, executed emergency changes, and documented runbooks and knowledge articles.

Technology areas

Observability

AppDynamics, Splunk, ELK/Kibana, ThousandEyes, GenEos, Wily Introscope

ITSM & Knowledge

ServiceNow, ServiceFirst, SharePoint, Confluence, Jira

SRE, Resiliency & Controls

Incident response, RCA, runbooks, service health, Cutover, SRP, PARP, DIRP, DCR, and risk story tracking

CI/CD & Automation

Jenkins, Chef, GoCD, Rundeck, Git

Integration & Messaging

REST, SOAP/WSDL, IBM MQ, Kafka, JSON, partner connectivity models

Platforms & Data

AWS, WebLogic, WebSphere, ActiveMQ, Unix/Linux, Windows Server, Oracle, Sybase, MongoDB, SQL Server

Education

M.S. Computer Science

Virginia International University, Fairfax, VA

B.Tech Electronics & Communication Engineering

Vignan's Institute of Information and Technology, India

Let's connect about SRE, API platform support, and reliable production operations.

Based in Whippany, NJ.