Kubernetes Cluster Health Monitoring and Failure Detection System

NTPL Digital Private Limited Cloud Infrastructure & Devops Operations
LocationRemote
#HiringActivily
#TopOpportunity

Project Objectives:

Develop a Kubernetes cluster health monitoring system that tracks node performance, pod health, and network conditions to ensure high availability and early failure detection.

Project Tasks:

Deploy Kubernetes cluster.

Install Prometheus Operator.

Configure kube-state-metrics.

Monitor pod lifecycle events.

Track node CPU/memory usage.

Configure alert rules for node failures.

Simulate pod crashes.

Analyze restart behavior.

Implement dashboard visualizations.

Document cluster observability architecture.

Educational Qualifications

B.TechB.EBCAMCA

Required Skills

Cloud Infrastructure & Deployment ManagementAlerting & Incident Management (Pagerduty, Opsgenie)Kubernetes Cluster ManagementMonitoring Tools (Prometheus, Grafana)Performance Analysis & Troubleshooting