system overview --:--:-- UTC

Darshan Jain

Darshan Jain
IMG_0042 ✳ subject: darshan-jain

role: loading…

Breaking infrastructure on purpose,
so production stays up when it matters.

chaos@darshan — boot sequence
experience
7+ yrs
backend & infrastructure
throughput
10K/min
events ingested at peak
infra cost
−$40K
saved per year on AWS
bug yield
−90%
after layered testing effort
[ 01 ]

Work

2026/01 → now · 9 mo

Senior Software Engineer

Red Hat · Performance & Scale · Pune, Maharashtra, IN

  • Maintainer of Krkn, a CNCF open-source chaos engineering framework for Kubernetes and OpenShift — driving roadmap, feature development, release management, architecture decisions, code reviews, and community growth
  • Owned resilience validation for Cross-Cluster Live Migration (CCLM), using chaos engineering to validate VM migration under real-world failure scenarios and establish reliability benchmarks, scalability guidance, and operational best practices for OpenShift Virtualization
  • Led the design and implementation of the Storage I/O Throttling chaos scenario — controlled IOPS and bandwidth degradation across cgroup v1/v2 environments for PVC-backed workloads, powered by a new deploy_io_throttle_pod helper in krkn-lib — with documentation and krkn-hub support shipped in parallel
  • Architected a next-generation pytest-based CI test framework v2 with ephemeral namespace isolation and automated reporting — including the memory-hog and cpu-hog v2 test suites — modernizing legacy validation workflows
  • Introduced event-driven chaos triggers — command-trigger Phase 1 wired through global krknctl inputs — so scenario execution can react to failure conditions the moment they surface
  • Hardened the SLO engine — parallelizing PromQL queries with ThreadPoolExecutor and eliminating duplicate SLO evaluation between per-scenario and final scoring
  • Contribute to the CNCF ecosystem as an LFX Mentor, guiding contributors on automation tooling and docs consistency — and presented Krkn at DevConf Pune 2026
2026/06 → 2026/08 · 3 mos

LFX Mentor — Krkn-Chaos

Cloud Native Computing Foundation · CNCF/LFX mentorship · 2026 term 2

  • Mentored an open-source contributor through the CNCF/LFX Mentorship Program on an automated documentation-synchronization bot for the Krkn-Chaos ecosystem
  • Guided the design of a GitHub Actions bot workflow — Python, Hugo/Docsy, GitHub automation, and LLM agents — that detects documentation-impacting changes and generates draft documentation PRs
  • Supported the mentee end-to-end: architecture discussions, problem-solving, open-source contribution practices, code reviews, and iterative improvement toward a practical shipped solution
  • project record ↗
2018/10 → 2025/12 · 7 yrs 3 mos

Senior Lead Software Engineer

HeapTrace Technology · healthcare infrastructure

  • Architected multi-tenant backend services in Spring Boot with role-based access controls
  • Built high-performance ingestion APIs using Falcon framework, handling 10K+ events/min
  • Designed AWS SQS pipeline for real-time, fault-tolerant patient vitals ingestion
  • Saved $40,000 annually by shifting EC2 workloads to Reserved Instances
  • Migrated jobs to Lambda, reducing costs by 75% and improving speed by 5x
  • Implemented Blue-Green deployments for 15+ ECS microservices
  • Architected HIPAA-compliant Engage app with OAuth2/Keycloak auth, Amazon Lexbot, Lambda, and Twilio API
  • Mentored 4+ junior developers; led testing strategy that raised coverage to 85% and reduced bugs by 90%
[ 02 ]

Lab

mcp — kubectl via chat

Kubernetes MCP Server

MCP server that lets AI agents inspect and manage Kubernetes clusters via natural language—useful when you want to troubleshoot or run kubectl-style tasks from a chat interface.

PythonKubernetesMCPAI
view on github →
rag — qdrant vector store

RAG Application with Qdrant

RAG app using the Qdrant vector database so you can ask questions over your own documents instead of searching manually. Good for internal docs or knowledge bases.

PythonQdrantRAGAI
view on github →
[ 03 ]

Stage

  • 2026

    DevConf Pune — Krkn: breaking clusters so yours don't fold

    Presented Krkn; gathered field feedback from IT professionals and contributed to roadmap and feature implementation.

    recorded
    live ✓
[ 04 ]

Profile

Senior Software Engineer at Red Hat on the Performance & Scale team — and maintainer of Krkn, a CNCF open-source chaos engineering framework for testing the resilience of Kubernetes, OpenShift, and OpenShift Virtualization.

My days are spent breaking clusters on purpose — from CCLM migration validation to Storage I/O throttling chaos — then writing the benchmarks and best practices that keep production boring. Before Red Hat, 7+ years at HeapTrace building backend systems with Java, Spring Boot, and Python: multi-tenant services, high-throughput APIs, HIPAA-compliant applications, and cost-saving infrastructure.

toolbox
Java 17Spring Boot 3PythonAWSKubernetesDockerRedisMySQLPostgreSQLTerraformArgoCDPrometheusGrafanaOAuth2Keycloak
[ 05 ]

Log

$ no posts yet — follow LinkedIn / X ↗