Skip to content

Cluster update times out while fetching compute node LaunchTemplates for queues #7519

Description

@inoue-hideki-a-opc

Required Info

  • AWS ParallelCluster version: 3.15.0
  • OS: Ubuntu 24.04 LTS
  • Region: us-east-1

Bug description and how to reproduce

When queue settings are changed and the cluster is updated with pcluster update-cluster, the process times out while the Chef run on the head node retrieves the Compute Node LaunchTemplates associated with the queues.

We believe this occurs because the current implementation assumes that processing all queues will complete within 30 seconds.

As a temporary workaround, we changed the timeout setting from 30 to 300 directly in the following file:

  • /etc/chef/cookbooks/aws-parallelcluster-platform/resources/fetch_dna_files.rb

Expected behavior

As the number of queues increases, simply extending the timeout does not provide a fundamental solution. The timeout should apply to the processing of each queue rather than to the processing of all queues together.

Related Issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions