Fault tolerant distributed training on Amazon EKS using NVRx
AWS Machine Learningen

Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.
This is a short summary published by AI Global Wire. The full article is owned and hosted by AWS Machine Learning — open it there to read it in full.
Read the full story at AWS Machine LearningRelated AI news
- Improving HCLS AI reasoning with open-source agent skillsAWS Machine Learning · September 16, 2026
- Optimizing agent system prompts with Amazon Bedrock AgentCoreAWS Machine Learning · September 16, 2026
- Build a serverless PII redaction pipeline with Amazon Bedrock Data AutomationAWS Machine Learning · September 16, 2026
- Optimizing cost and latency with Amazon Bedrock prompt cachingAWS Machine Learning · September 15, 2026
- Build an AI-powered product tagging system with Amazon SageMaker serverless model customizationAWS Machine Learning · September 15, 2026
- Announcing instance preference lists for Amazon SageMaker AI training jobsAWS Machine Learning · September 15, 2026