Nagaraju Nampally

Nagaraju Nampally

Platform / Database Reliability Engineer (Lead)

9+ years specializing in PostgreSQL, cloud migrations, Kubernetes (EKS), and SRE automation

Passionate about building reliable, highly available platform infrastructure, automating cloud environments, and optimizing database workloads. Expert in multi-cloud operations, containerized services (EKS/GKE), GitOps, and platform toolings.

postgresql_console — nampally_db
Connected to PostgreSQL database "nampally_db" (v16.1)
Type a query or click the shortcut buttons below to interact.
nampally_db=> SELECT * FROM overview;
+----------------------+----------------------------------------------------------+ | specialization | focus_areas | +----------------------+----------------------------------------------------------+ | Lead DBRE & SRE | PostgreSQL, Kubernetes (EKS), IaC (Terraform), CI/CD | +----------------------+----------------------------------------------------------+ (1 row)
nampally_db=>
Queries:

About Me

Results-driven Platform & Database Reliability Engineer with 9+ years of experience specializing in PostgreSQL, cloud-native administration, database performance tuning, SRE automation, and multi-cloud architectures.

Database Tuning & Performance

Optimized indexes, query strategies, and parameters to achieve 300% throughput increases across high-concurrency systems.

Cloud-Native Platform Engineering

Engineered self-healing EKS and GKE Kubernetes cluster infrastructures, achieving 99.99% uptime and slashing cloud spend by 40%.

DevSecOps & Platform Automation

Implemented security scanning (SAST/DAST), ArgoCD GitOps pipelines, secret rotation, and IaC automation via Terraform and Ansible.

Dedicated to platform health, rigorous SRE observability (SLOs/SLIs), and removing operational toil. I build stable infrastructure pipelines that accelerate software delivery while maintaining rock-solid systems integrity.

9+
Years Experience
99.99%
Uptime Led
30%
Toil Reduced

Professional Experience

Platform / Database Reliability Engineer (Lead)

May 2025 – Present

ADP Inc. – Pasadena, CA

  • Lead platform reliability and database automation initiatives on AWS (EKS, EC2, RDS, Lambda, IAM, S3).
  • Design and build end-to-end observability pipelines using CloudWatch, Grafana, and Splunk, improving MTTR.
  • Automate platform lifecycle tasks (backups, failover validation, auto-scaling) using Terraform and Ansible IaC.
  • Implement DevSecOps practices including container scanning (SAST/DAST), secret management, and policy-as-code.
  • Orchestrate incident response and perform root cause analysis (RCA) to drive structural uptime improvements.

Senior Platform / Site Reliability Engineer

July 2021 – April 2025

Workday Inc. – Pleasanton, CA

  • Engineered and operated highly available PostgreSQL clusters and led zero-downtime MySQL migrations.
  • Managed Kubernetes (EKS/GKE) deployments, service meshes, scaling configurations, and Helm charts.
  • Built CI/CD pipelines with GitHub Actions and GitOps (ArgoCD), reducing release cycles from 3 months to 7 days.
  • Created Python-based APIs (Flask) for platform observability and developer self-service tooling.
  • Enabled MLOps pipelines using AWS SageMaker and MLflow, accelerating AI/ML team delivery velocities.
  • Participated in 24/7 on-call rotations, establishing automated self-healing mechanisms to resolve incidents.

Cloud Platform Engineer / AWS Solutions Architect (Lead)

Aug 2018 – July 2021

Synapsis Inc. (T-Mobile, Deloitte)

  • Migrated legacy SQL Server and Oracle databases to high-performance PostgreSQL on AWS RDS.
  • Designed multi-AZ and cross-region high availability/disaster recovery (HA/DR) architectures.
  • Created centralized monitoring and alerting dashboards using ELK, Prometheus, and CloudWatch.
  • Optimized AWS cloud resource utilization, generating substantial cost savings through rightsizing.

Cloud Infrastructure & Automation Engineer

May 2016 – Aug 2018

PVR America & Global Data Mart

  • Administered production Linux servers, enforcing system hardening and network security policies.
  • Automated database maintenance and recurring system administration using Python and Shell scripting.
  • Built robust data pipeline integrations and supported ETL/ELT data migration workflows.

Technical Skills

Database Technologies

PostgreSQL Performance Tuning Database Design MySQL Redis MongoDB Redshift Oracle

Cloud & Infrastructure

AWS EKS / RDS / EC2 GCP GKE Azure App Services VPC Networking IAM Security Cloud Storage

DevOps & GitOps Tools

Kubernetes Docker Helm Charts Terraform Ansible & Chef ArgoCD (GitOps) GitHub Actions

Observability & SRE

Prometheus Grafana Splunk ELK Stack AlertManager SLO / SLI Tracking

Programming Languages

Python (Django/Flask) Bash Scripting SQL / PL-pgSQL Go Language

Methodologies

MLOps (SageMaker) DevSecOps (SAST/DAST) Policy-as-Code Zero-Downtime Migration HA/DR Failover

Featured Projects

Memzent.AI

An intelligent semantic proxy and memory layer for autonomous AI agents. Features multi-layer semantic caching and GPU avoidance to slash LLM costs/latency by 80-90% with mTLS security, prompt audits, and hardware JWT validation.

AI Agents Semantic Caching Python gRPC JWT
Visit Repository

IncognitoJSON

A privacy-focused client-side JSON formatting and validation tool built with modern web technologies. Features secure offline data processing, syntax highlighting, and formatting without any server-side logs or storage.

JavaScript HTML5 CSS3 JSON Processing
Visit Website

OpsyLux

A SaaS event hall booking platform. Built with a scalable multi-tenant architecture, incorporating strict tenant database isolation, JWT-based security protocols, containerized deployments, and Prometheus observability.

Django PostgreSQL Kubernetes Prometheus
Visit Website

Other Engineering Solutions

Multi-Cloud Database Migration Suite

Automated scripts and schemas to migrate database instances to AWS RDS and Azure SQL, cutting license fees and operations work.

PostgreSQL AWS DMS Terraform

SRE Automated Platform Observability

Custom metrics collection agent integrating Prometheus and Splunk to notify on cluster health issues, dropping response times.

Grafana Prometheus Python

High-Performance Indexing & Query Tuning

Advanced indexing and query rewrite tools optimizing latency by 300% under high read-write churn applications.

PostgreSQL Query Rewrite DB Tuning

Get In Touch

Ready to discuss infrastructure challenges, database operations, or SRE automation? I'd love to hear from you.