Skip to content

cluster update fails in 3.10.0, 3.9.3 #6339

Description

@snemir2

If you have an active AWS support contract, please open a case with AWS Premium Support team using the below documentation to report the issue:
https://docs.aws.amazon.com/awssupport/latest/user/case-management.html

Before submitting a new issue, please search through open GitHub Issues and check out the troubleshooting documentation.

Please make sure to add the following data in order to facilitate the root cause detection.

Required Info:

  • AWS ParallelCluster version [e.g. 3.1.1]: 3.10.0
  • Full cluster configuration without any credentials or personal data.
  • Cluster name: A2AiClustertesting
  • Output of pcluster describe-cluster command.
 pcluster describe-cluster -n A2AiClustertesting -r us-east-2
{
  "creationTime": "2024-07-09T18:49:23.141Z",
  "headNode": {
    "launchTime": "2024-07-09T18:53:32.000Z",
    "instanceId": "i-0976556062851f6ca",
    "instanceType": "r6i.xlarge",
    "state": "running",
    "privateIpAddress": "10.2.46.69"
  },
  "version": "3.10.0",
  "clusterConfiguration": {
    "url": "https://parallelcluster-97a8b56da16cbe1e-v1-do-not-delete.s3.us-east-2.amazonaws.com/parallelcluster/3.10.0/clusters/a2aiclustertesting-3vtysbdo94ne39ev/configs/cluster-config.yaml?versionId=rhu_16ixsYqMF2ctGyWRVVtNGS.jwCQI&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIAZQUXECJHGFFWWCDI%2F20240709%2Fus-east-2%2Fs3%2Faws4_request&X-Amz-Date=20240709T212526Z&X-Amz-Expires=3600&X-Amz-SignedHeaders=host&X-Amz-Security-Token=IQoJb3JpZ2luX2VjEG0aCXVzLWVhc3QtMiJHMEUCIFlopqMIbJ6IffFlaCwfCvGgUh0RIeHnnlInHBRnLECNAiEAnl07d3CZ39jLquycKjIjGyDcuvvzOR%2FiJ7vAVG%2FBRDAq%2FgIINxAAGgw2NTQyMjU3MDc1OTgiDNQTAke7TLCxmdDjFirbAlrYusjjQ3oD4XjjeNPyzpzxeX6as8JfiomXPRwmzsHOJl7ttg11miKNZ1h4h%2Fgt2MN%2FVucaVJoc%2BnWfHiXHQ8PTfWqjisZ698iw2QrMzLYzatufZSuwpfumz93eH1E8UCtNctjCvUdIqsr6vwTXFKoPqXhKm5KgZ5pfgqK381VNQmFP1xxPqnflpyL0pRnIRBC76XWdaD1zNAZluzp0Zxce75MiXjPT1NPqqu%2Fcux3VSTHgvPbuJfF2yri5pfRpp7n7KiLHgBus8OAfM%2FEwFMLvtnNPP61Hk%2BU0YvWZvuuXF6lLisxqxw4wZYNB0zR7zF3GecXDvuW4ZS%2Bapdme8hzOCk4xh4XG271G1p6Ch%2FG%2BvIWF4roQGgJBu3mOWrOEERzihvgeEDCZUsyIhJnUrJjSPuYsfAlf7aDDvEgru4sW2tKjCsShGpph%2F3cQa2hw4Y1k0DaUSCHPVZ5GMMXVtrQGOqcBeG9WoHb5rdd%2FG0uUI3pfUDWLVFC%2FyswoY22gb0Rkk8GIb3bBSm9SrYZOmxXw5lz%2FOP8X446KfcLMzInE2WeSZ9cijK5RT%2FAywuCQm4yXCglbra%2B0OG0r%2BWc%2BX0MkFDrtepKJRKeMH7pesvzqm8MWkqWpUUC59r9u%2BKa58HZQ6jjx0Icl00MUxa17OqYzQ0vUZqADUxggW9QddoJcdcpKLSlKUkEfjSY%3D&X-Amz-Signature=63e0e16188b1f9e6c696e5b980610f33e775c3fd801df5c1b4d618362eba722a"
  },
  "tags": [
    {
      "value": "branch/release-v4.0",
      "key": "A2AI:a2ai-cloud-version"
    },
    {
      "value": "mig8KP4B19EMB",
      "key": "map-migrated"
    },
    {
      "value": "3.10.0",
      "key": "parallelcluster:version"
    },
    {
      "value": "A2AiClustertesting",
      "key": "parallelcluster:cluster-name"
    },
    {
      "value": "sergey",
      "key": "A2AI:creator"
    },
    {
      "value": "dev",
      "key": "A2AI:a2ai-cloud-env"
    }
  ],
  "cloudFormationStackStatus": "UPDATE_ROLLBACK_COMPLETE",
  "clusterName": "A2AiClustertesting",
  "computeFleetStatus": "RUNNING",
  "cloudformationStackArn": "arn:aws:cloudformation:us-east-2:654225707598:stack/A2AiClustertesting/f531bdf0-3e23-11ef-997c-06835d7b2d0f",
  "lastUpdatedTime": "2024-07-09T19:32:18.912Z",
  "region": "us-east-2",
  "clusterStatus": "UPDATE_FAILED",
  "scheduler": {
    "type": "slurm"
  }
}
  • [Optional] Arn of the cluster CloudFormation main stack:

Bug description and how to reproduce:
A clear and concise description of what the bug is and the steps to reproduce the behavior.

Cluster repeatedly fails to update and from cloud-formation point of view goes to "rollback complete" . (custom routines do not appear even to get called)

If you are reporting issues about scaling or job failure:
We cannot work on issues without proper logs. We STRONGLY recommend following this guide and attach the complete cluster log archive with the ticket.

For issues with Slurm scheduler, please attach the following logs:

  • From Head node: /var/log/parallelcluster/clustermgtd, /var/log/parallelcluster/clusterstatusmgtd (if version >= 3.2.0), /var/log/parallelcluster/slurm_resume.log, /var/log/parallelcluster/slurm_suspend.log, /var/log/parallelcluster/slurm_fleet_status_manager.log (if version >= 3.2.0) and/var/log/slurmctld.log.
  • From Compute node: /var/log/parallelcluster/computemgtd.log and /var/log/slurmd.log.

If you are reporting issues about cluster creation failure or node failure:

If the cluster fails creation, please re-execute create-cluster action using --rollback-on-failure false option.

We cannot work on issues without proper logs. We STRONGLY recommend following this guide and attach the complete cluster log archive with the ticket.

Please be sure to attach the following logs:

  • From Head node: /var/log/cloud-init.log, /var/log/cfn-init.log and /var/log/chef-client.log
    (attached)
  • From Compute node: /var/log/cloud-init-output.log.
    NA
    logs.tgz <-headnode logs

Additional context:
Any other context about the problem. E.g.:

  • CLI logs: ~/.parallelcluster/pcluster-cli.log
  • Custom bootstrap scripts, if any
  • Screenshots, if useful.
    image

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions