Muhammed Ajas
I build reliable systems, automate the difficult parts, and make production engineering better.
I'm a Senior Site Reliability Engineer and Cloud Architect working across cloud infrastructure, Kubernetes, automation, observability, security, and production reliability.
My work focuses on turning complex operational problems into reliable engineering solutions. When something repeatedly requires manual intervention, I look for a way to automate it. When a production failure is difficult to understand, I look for better diagnostics and better evidence.
Currently at IBM, building and improving infrastructure and operational capabilities for Cognos Analytics on Cloud—managing 300+ Kubernetes clusters serving 500+ enterprise customers.
$ whoami
> Senior Site Reliability Engineer & Multi-Cloud Architect
$ systemctl status availability
● active (running) — 99.99% multi-tenant uptime
$ cat /proc/expertise
> Kubernetes • AWS • IBM Cloud • Automation • Security

WHAT I DO
Site Reliability Engineering
Improving the reliability, observability, maintainability, and recoverability of production systems.
Cloud Architecture
Designing cloud infrastructure with a focus on security, reliability, maintainability, and operational simplicity.
Kubernetes
Working with Kubernetes platforms, workloads, networking, configuration, troubleshooting, Helm, and operational lifecycle management.
Automation
Turning repetitive operational activities into reliable workflows using Python, Bash, APIs, Kubernetes, and infrastructure tooling.
Production Engineering
Investigating complex failures across application, infrastructure, database, networking, authentication, and platform layers.
Incident Response
Leading technical investigations, coordinating recovery, communicating during incidents, and driving issues toward permanent remediation.
Platform Engineering
Building internal tools and self-service capabilities that make infrastructure operations easier and more consistent.
Security Automation
Automating credentials, secrets, certificates, and security-related operational processes.
WORK EXPERIENCE
Senior Site Reliability Engineer
Leading reliability engineering for Kubernetes infrastructure. Managing 300+ production clusters serving 500+ customers across multiple cloud platforms.
The role combines hands-on engineering with operational ownership—understanding how systems behave in production and continuously improving how they are operated.
- •Designed RootCause Helper - automated forensic diagnostics platform cutting MTTR by 40%
- •Architected Customer Self-Service Platform replacing manual onboarding for 500+ customers
- •Reduced monthly cloud costs by $19K through Kubernetes optimization and FinOps automation
- •Served as Incident Commander for 10+ Sev-1 incidents and led disaster recovery operations
- •Led and mentored team of 11 Site Reliability Engineers
Cloud Architect / Site Reliability Engineer
Led cloud infrastructure discussions and migrated enterprise workloads from on-premises to AWS with zero downtime.
- •Designed secure AWS networking using VPC Peering, Transit Gateway, and Direct Connect
- •Implemented AWS IAM governance, GuardDuty monitoring, and Secrets Manager integration
- •Built automation using AWS Lambda and Python for operational workflows
Cloud Engineer
Managed enterprise hybrid cloud operations including OS patching, backup, and infrastructure maintenance. Configured CloudWatch monitoring and proactive alerting.
System Engineer
Served as primary technical contact for global enterprise customers, resolving complex incidents while meeting SLA commitments using ITIL best practices.

PROJECTS
RootCause Helper
When production fails, the evidence should already be there.
An automated forensic diagnostics platform that preserves application, Kubernetes, database, infrastructure, and network logs before pod or node recovery—supporting approximately 15 production incidents per month and materially cutting Mean Time to Recovery (MTTR) by 40%.
The platform automatically gathers and preserves important diagnostic information before recovery actions remove valuable evidence. Traditional troubleshooting often starts with "Let's see what logs we still have." RootCause Helper changes that to: "The evidence has already been collected. Now let's understand what it tells us."
Customer Self-Service Platform
Turning operational requests into engineering workflows.
Flask-based platform that replaced manual email-driven onboarding by automating SSL uploads, custom URLs, SMTP configuration, contact management, and controlled cluster restarts for 500+ enterprise customers.
Instead of engineers manually processing similar requests repeatedly, the platform provides a structured workflow for common operational activities. This improves consistency, reduces manual configuration errors, and makes routine environment management easier to execute.
FinOps Automation Pipeline
Automated cost optimization pipeline through Kubernetes resource right-sizing, CPU/memory optimization, automated orphaned-storage cleanup, and improved cluster decommission workflows, achieving $19K+ monthly savings.
The goal was not simply to reduce resource consumption—it was to make resource efficiency part of normal engineering operations.
Enterprise Security Automation
Automated enterprise security and certificate compliance using HashiCorp Vault APIs for password rotation, audit automation, legacy-credential detection, and full SSL/Kubernetes Secret lifecycle management, preventing certificate-expiration incidents.
Security controls should have an operational lifecycle just like application infrastructure does.
ENGINEERING TOOLBOX
Cloud Platforms
- ▸ AWS (EKS, EC2, Lambda, IAM)
- ▸ IBM Cloud
- ▸ Microsoft Azure
- ▸ Google Cloud Platform
Container & Orchestration
- ▸ Kubernetes
- ▸ Docker & Podman
- ▸ Helm
- ▸ Container Security
IaC & Automation
- ▸ Terraform
- ▸ Ansible
- ▸ Jenkins & GitHub Actions
- ▸ Python & Bash
Observability
- ▸ Instana
- ▸ Sysdig
- ▸ IBM Cloud Logs
- ▸ AWS CloudWatch & Splunk
Security & Compliance
- ▸ HashiCorp Vault
- ▸ AWS Secrets Manager
- ▸ IAM & GuardDuty
- ▸ SSL/TLS Management
SRE & Operations
- ▸ Incident Management
- ▸ Disaster Recovery
- ▸ PagerDuty & ServiceNow
- ▸ ITIL Best Practices
CERTIFICATIONS
AWS Certified Solutions Architect
Advanced certification demonstrating expertise in designing distributed systems on AWS.
AWS Certified Solutions Architect
Foundational certification in AWS architecture and best practices.
ITIL Foundation
Certification in IT service management best practices and frameworks.
COMMUNITY
Progressive Techies Kerala
Core group member delivering technical sessions on Kubernetes, SRE, Cloud Automation, and DevOps best practices for engineering teams and technical communities. I enjoy sharing practical engineering experiences and turning lessons from production into discussions that other engineers can learn from.
DevOps Malayalam Tech Community
Active moderator fostering technical discussions and knowledge sharing in the regional DevOps community.
NGO Career Guidance Program
Mentor providing career guidance and technical mentorship to aspiring engineers. I enjoy helping early-career engineers understand how to move from learning technology to building practical engineering skills.
Panayikulam Public Library
Secretary supporting community education and literacy initiatives.
Toastmasters Club
Active member developing public speaking and leadership skills. Engineering is only part of the job—being able to explain an idea clearly and communicate during difficult situations is equally important.

WRITING & IDEAS
ENGINEERING PHILOSOPHY
Reliability is engineered, not monitored into existence
Monitoring can tell you that something is wrong. Reliability engineering asks: Why did it happen? How do we recover safely? How do we prevent it from happening again?
Automate the operational knowledge
If engineers repeatedly perform the same investigation, recovery process, or configuration task, there is an opportunity to turn that knowledge into software.
Preserve evidence
A recovery that destroys the evidence needed for diagnosis can make the next failure harder to understand. Good operational systems should make investigation easier, not harder.
Make complexity understandable
Infrastructure can become complicated very quickly. Good engineering is not only about building sophisticated systems—it is also about creating the tools, documentation, automation, and observability that allow humans to understand those systems.
Automate with purpose
Automation isn't about removing humans from every process. It is about removing repetitive work so engineers can spend their time on problems that actually require engineering judgment.
Keep learning
Cloud platforms evolve. Kubernetes evolves. Automation evolves. AI is changing how we think about infrastructure operations. The tools will change. The willingness to learn should not.
BEYOND THE SCREEN
"There is more to me than technology."
Engineering is a big part of my life, but curiosity doesn't stop when I close my laptop. I enjoy travelling, discovering unfamiliar places, challenging myself with new experiences, playing sports, and exploring creative expression. Many of the things I enjoy outside technology have something in common with engineering: curiosity, preparation, adaptability, persistence, and the willingness to try something new.

Travel Enthusiast
Exploring diverse landscapes and cultures across the globe. For me, travelling isn't simply about visiting destinations—it's about experiencing different cultures, meeting people, and seeing the world from a different point of view.

Adventure Sports
Pushing boundaries through high-adrenaline outdoor pursuits including skydiving, river rafting, kayaking, and mountaineering. I learned to swim at age 30 and later challenged myself to cross the Periyar river. Being a beginner can be uncomfortable—it can also be one of the most rewarding places to be.

Sports & Creative Expression
Competitive badminton player, Valasery Kayaking Club member, and exploring creative expression through acting and public speaking. There is something surprisingly similar between engineering and communication: take something complicated, understand it deeply, then make it simple for someone else.
LET'S CONNECT
Interested in discussing SRE practices, cloud architecture, Kubernetes, platform engineering, infrastructure automation, or AI & agentic reliability? Whether you're working through a difficult production problem or simply want to exchange engineering perspectives, I'd be happy to connect.