Required Info
- AWS ParallelCluster version: 3.15.0
- OS: Ubuntu 24.04 LTS
- Region: us-east-1
Bug description and how to reproduce
When queue settings are changed and the cluster is updated with pcluster update-cluster, the process times out while the Chef run on the head node retrieves the Compute Node LaunchTemplates associated with the queues.
We believe this occurs because the current implementation assumes that processing all queues will complete within 30 seconds.
As a temporary workaround, we changed the timeout setting from 30 to 300 directly in the following file:
/etc/chef/cookbooks/aws-parallelcluster-platform/resources/fetch_dna_files.rb
Expected behavior
As the number of queues increases, simply extending the timeout does not provide a fundamental solution. The timeout should apply to the processing of each queue rather than to the processing of all queues together.
Related Issues
Required Info
Bug description and how to reproduce
When queue settings are changed and the cluster is updated with
pcluster update-cluster, the process times out while the Chef run on the head node retrieves the Compute Node LaunchTemplates associated with the queues.We believe this occurs because the current implementation assumes that processing all queues will complete within 30 seconds.
As a temporary workaround, we changed the timeout setting from
30to300directly in the following file:/etc/chef/cookbooks/aws-parallelcluster-platform/resources/fetch_dna_files.rbExpected behavior
As the number of queues increases, simply extending the timeout does not provide a fundamental solution. The timeout should apply to the processing of each queue rather than to the processing of all queues together.
Related Issues