Hi all, I have encountered a set of issues involving promotion/demotion in dqlite that appear to impact fault tolerance. The issue reported here is essentially: a promotion will fail whenever the portion of the log required to catch-up the potential-promotee exceeds MAX_PAYLOAD_LEN, but there is no handling of this requirement on the promoter's side. This has the consequence of making it possible for all nodes except for the leader and voters, including stand-bys, to be unpromotable-to-voter without warning.
I encountered this issue when attempting to administer clusters with sizes larger than 50 nodes (average log entry size is a factor in how large the unpromotability window is and scales significantly with the number of nodes at least in the users of dqlite I have been testing). I've attached a reproducer below which demonstrates this problem with a synthetically-inflated raft log, both in a scenario that shows an unpromotable spare and in one where this problem makes all nodes including standbys unpromotable except for active voters such that any failure cannot promote a voter. I will talk a little bit about how this is showing up in our scaling tests after explaining the mechanism in case there is some concern that the reproducer is producing artificially large log files.
uv_recv.c defines limit macros:
|
#define MAX_HEADER_LEN (4 * 1024 * 1024) /* 4 MiB */ |
|
#define MAX_PAYLOAD_LEN (500 * 1024 * 1024) /* 500 MiB */ |
in case the message exceeds MAX_PAYLOAD_LEN, the receiver aborts the communication:
|
} else if (s->payload.len > MAX_PAYLOAD_LEN) { |
|
tracef("message payload too long: %" PRIu64 " bytes", |
|
(uint64_t)s->payload.len); |
|
if (s->message.type == RAFT_IO_APPEND_ENTRIES) { |
|
raft_free(s->message.append_entries.entries); |
|
s->message.append_entries.entries = NULL; |
|
s->message.append_entries.n_entries = 0; |
|
} |
|
goto abort; |
|
} |
The sender places no limit on the catch-up size and does not otherwise handle this case:
|
/* TODO: implement a limit to the total size of the entries being sent |
|
*/ |
|
rv = logAcquire(r->log, next_index, &args->entries, &args->n_entries); |
|
if (rv != 0) { |
|
goto err; |
|
} |
and then sends the potentially-too-large catch-up to the server:
|
req->raft = r; |
|
req->index = args->prev_log_index + 1; |
|
req->entries = args->entries; |
|
req->n = args->n_entries; |
|
req->server_id = server->id; |
|
|
|
req->send.data = req; |
|
rv = r->io->send(r->io, &req->send, &message, sendAppendEntriesCb); |
|
if (rv != 0) { |
|
goto err_after_req_alloc; |
|
} |
The consequence is that the promotion hangs for max_catch_up_round_duration (50s), during which time role adjustment is frozen and the leader's configuration change slot is held.
This also affects stand-bys: catch-up does not prevent them from being promoted to stand-by from spare, and as stand-bys they silently fail to catch up due to the 500MiB limit such that they cannot be promoted to voter.s.
This is a window of unpromotability that lasts until the last seen sequence number of the unpromotable node is compacted out from the logs, but the window can reopen again any time a node is demoted (for that node).
In real-world scenarios, catch-up logs do reach significantly more than 500MiB. In a microcloud/microovn deployment at 100 nodes, I observed an average log entry size of about 240KiB per entry at the full deployment scale. This issue interacts with a go-dqlite bug which causes every new joiner to self-promote to create a situation where a very large proportion of the cluster becomes unpromotable during an overlapping period of time, during which fault tolerance is correspondingly seriously impaired. In a large enough cluster, any demotion starts a clock essentially: after 500MiB of logs have been written, the node becomes unpromotable until compaction removes the last seen seqno from the log file.
This issue would also affect any situations where the snapshot file itself is larger than 500MiB, although I have not reproduced that and don't know if it's possible - just that there's nothing I can see preventing it. If it is possible it would cover the inverse of the current window essentially, and might make promotion impossible generally except for a small window.
I've attached reproducers that demonstrate this problem synthetically against the go-dqlite demo. Note that this has artificially inflated log entry sizes and that in general never-voter spares are probably not vulnerable to this until cluster sizes of 200+ members (at least in microovn/microcloud clusters), but demoted-voter spares/standbys are starting at a much lower level.
dqlite-issue.tar.gz
Hi all, I have encountered a set of issues involving promotion/demotion in dqlite that appear to impact fault tolerance. The issue reported here is essentially: a promotion will fail whenever the portion of the log required to catch-up the potential-promotee exceeds MAX_PAYLOAD_LEN, but there is no handling of this requirement on the promoter's side. This has the consequence of making it possible for all nodes except for the leader and voters, including stand-bys, to be unpromotable-to-voter without warning.
I encountered this issue when attempting to administer clusters with sizes larger than 50 nodes (average log entry size is a factor in how large the unpromotability window is and scales significantly with the number of nodes at least in the users of dqlite I have been testing). I've attached a reproducer below which demonstrates this problem with a synthetically-inflated raft log, both in a scenario that shows an unpromotable spare and in one where this problem makes all nodes including standbys unpromotable except for active voters such that any failure cannot promote a voter. I will talk a little bit about how this is showing up in our scaling tests after explaining the mechanism in case there is some concern that the reproducer is producing artificially large log files.
uv_recv.c defines limit macros:
dqlite/src/raft/uv_recv.c
Lines 13 to 14 in 9055088
in case the message exceeds MAX_PAYLOAD_LEN, the receiver aborts the communication:
dqlite/src/raft/uv_recv.c
Lines 283 to 292 in 9055088
The sender places no limit on the catch-up size and does not otherwise handle this case:
dqlite/src/raft/replication.c
Lines 75 to 80 in 9055088
and then sends the potentially-too-large catch-up to the server:
dqlite/src/raft/replication.c
Lines 106 to 116 in 9055088
The consequence is that the promotion hangs for max_catch_up_round_duration (50s), during which time role adjustment is frozen and the leader's configuration change slot is held.
This also affects stand-bys: catch-up does not prevent them from being promoted to stand-by from spare, and as stand-bys they silently fail to catch up due to the 500MiB limit such that they cannot be promoted to voter.s.
This is a window of unpromotability that lasts until the last seen sequence number of the unpromotable node is compacted out from the logs, but the window can reopen again any time a node is demoted (for that node).
In real-world scenarios, catch-up logs do reach significantly more than 500MiB. In a microcloud/microovn deployment at 100 nodes, I observed an average log entry size of about 240KiB per entry at the full deployment scale. This issue interacts with a go-dqlite bug which causes every new joiner to self-promote to create a situation where a very large proportion of the cluster becomes unpromotable during an overlapping period of time, during which fault tolerance is correspondingly seriously impaired. In a large enough cluster, any demotion starts a clock essentially: after 500MiB of logs have been written, the node becomes unpromotable until compaction removes the last seen seqno from the log file.
This issue would also affect any situations where the snapshot file itself is larger than 500MiB, although I have not reproduced that and don't know if it's possible - just that there's nothing I can see preventing it. If it is possible it would cover the inverse of the current window essentially, and might make promotion impossible generally except for a small window.
I've attached reproducers that demonstrate this problem synthetically against the go-dqlite demo. Note that this has artificially inflated log entry sizes and that in general never-voter spares are probably not vulnerable to this until cluster sizes of 200+ members (at least in microovn/microcloud clusters), but demoted-voter spares/standbys are starting at a much lower level.
dqlite-issue.tar.gz