Repository navigation
fix(cluster): CLUSTER MIGRATE 真的把键搬走,并且只在目标确认后才删源键 #92
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change | ||||
|---|---|---|---|---|---|---|
|
|
@@ -211,8 +211,18 @@ RESTORE <key> <ttl> <serialized-value> | |||||
|
|
||||||
| - 从 `CLUSTER MIGRATE` 流程接收已序列化的 `CacheObject` | ||||||
| - `ttl` 单位为毫秒;0 表示永不过期 | ||||||
| - 配套序列化由 `CacheObject::serialize()` 提供 | ||||||
| - 错误返回:`-ERR invalid TTL`(ttl 非法)/ `-ERR invalid serialized data for <TYPE>`(反序列化失败)/ `-BUSYKEY Target key name already exists`(key 已存在且未带 REPLACE) | ||||||
| - 配套序列化由 `CacheObject::serialize()` 提供:一行类型标签 + 若干条 | ||||||
| `<字节数>\n<原始字节>` 记录,因此成员/字段里含 `\n`、`\r` 不会再被截断 | ||||||
| - 注意:总线本身用裸 `\xC0` 字节分隔参数且没有转义,所以**含 0xC0 的 value 在复制/迁移 | ||||||
| 时仍会在总线层被切断**(这是与载荷框架无关的另一处缺陷,登记在案待修) | ||||||
| - 载荷任何一帧不完整都算失败(不再"读到哪算哪"交出半个对象) | ||||||
| - 错误返回:`-ERR invalid TTL`(ttl 非法)/ `-ERR Invalid or malformed serialized payload` | ||||||
| (反序列化失败)/ `-BUSYKEY Target key name already exists`(key 已存在且未带 REPLACE) | ||||||
|
|
||||||
| `CLUSTER MIGRATE host port key destination-db timeout [REPLACE]` 走的是"发到目标 + 等它回复": | ||||||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
The documented MIGRATE syntax includes destination-db, but the handler implements only
Suggested change
Prompt to fix with AI
|
||||||
| 只有目标回 `+OK` 之后才删源键(默认语义是移动,不是复制)。超时、连不上、或目标回了 | ||||||
| `-BUSYKEY` 之类的错误时,**源键保持不动**,错误转给客户端。总线的请求/回复用 | ||||||
| `CCREQ <id> <RESP 命令>` / `CCRESP <id> <RESP 回复>` 两种标记,帧结构本身没有变。 | ||||||
|
|
||||||
| ## 15. 错误码 | ||||||
|
|
||||||
|
|
||||||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -14,6 +14,7 @@ | |
| #include <chrono> | ||
| #include <thread> | ||
| #include <atomic> | ||
| #include <condition_variable> | ||
|
|
||
| namespace cc_server { | ||
|
|
||
|
|
@@ -109,6 +110,24 @@ class ClusterConnection { | |
| // 向节点发送 RESP 命令(用于 MIGRATE 等场景) | ||
| bool send_command_to_node(const std::string& node_name, const std::vector<std::string>& args); | ||
|
|
||
| /** | ||
| * @brief 给节点发一条命令,并等它的 RESP 回复(带超时) | ||
| * | ||
| * CLUSTER MIGRATE 需要这个:只有确认目标节点收下(+OK)才能删源键。总线原本 | ||
| * 只有 kRepData 的单向推送、没有请求/回复关联,所以发送方永远不知道对端是 | ||
| * 接受了还是回了 -BUSYKEY —— 在那个前提下"补上删源键"等于可能把数据删没, | ||
| * 比留下重复键更糟。 | ||
| * | ||
| * 帧格式不动(header 保持原样):请求把命令行前加一个 "CCREQ <id> ",回复用 | ||
| * "CCRESP <id> ",两者仍是普通的 kRepData 参数,所以老的结构体长度测试不受影响。 | ||
| * | ||
| * @param timeout_ms 最长等待;到点返回 false,调用方因此不会去删源键 | ||
| * @return 拿到回复为 true(回复本身可能是错误回复,语义由调用方判断) | ||
| */ | ||
| bool send_command_and_wait(const std::string& node_name, | ||
| const std::vector<std::string>& args, | ||
| int timeout_ms, std::string& reply); | ||
|
|
||
| // 向节点发送原始字符串数据(用于复制命令推送) | ||
| bool send_raw_to_node(const std::string& node_name, const std::string& data); | ||
|
|
||
|
|
@@ -172,6 +191,19 @@ class ClusterConnection { | |
| std::unordered_map<int, Channel*> link_channels_; | ||
| std::mutex channel_mutex_; // 保护 link_channels_ | ||
|
|
||
| // 总线上的请求/回复关联。目前唯一的使用者是 CLUSTER MIGRATE。 | ||
| struct PendingBusReply { | ||
| std::string reply; | ||
| bool done = false; | ||
| int64_t deadline_ms = 0; // 回复比超时晚到时,这条会被下一个请求顺手清掉 | ||
| }; | ||
| // 收到 "CCRESP <id> ..." 时把回复交给等待者;id 不认识(已超时)就直接丢 | ||
| void deliver_command_reply(uint64_t request_id, const std::string& reply); | ||
| std::mutex pending_replies_mutex_; | ||
| std::condition_variable pending_replies_cv_; | ||
| std::unordered_map<uint64_t, PendingBusReply> pending_replies_; | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
A +OK from any accepted cluster link can satisfy this request because pending_replies_ is keyed only by request ID. A peer can guess the monotonic ID and forge CCRESP, causing MIGRATE to delete the source key even when the target never stored it. Prompt to fix with AI
|
||
| std::atomic<uint64_t> next_request_id_{1}; | ||
|
|
||
| NodeCallback node_connected_callback_; | ||
| NodeCallback node_disconnected_callback_; | ||
| ClusterLink::MsgCallback msg_callback_; | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -863,12 +863,35 @@ std::string ClusterCommand::handleMigrate(const std::vector<std::string>& args) | |
| restore_args.push_back("REPLACE"); | ||
| } | ||
|
|
||
| // 发送 RESTORE 命令到目标节点 | ||
| if (!conn->send_command_to_node(target_name, restore_args)) { | ||
| return RespEncoder::encode_error("ERR failed to send data to target node"); | ||
| } | ||
|
|
||
| LOG_INFO(CLUSTER, "MIGRATE completed: key=%s -> %s:%d", key.c_str(), host.c_str(), port); | ||
| // 把 RESTORE 发过去,并**等目标节点的回复**。 | ||
| // | ||
| // 两处原来都不对: | ||
| // 1. 以前用 send_command_to_node(),它只告诉你"发出去了没有",而回复被 | ||
| // 接收端丢弃 —— 于是这条命令其实从来没把数据搬走过(RESTORE 的参数 | ||
| // 被拆成多个总线参数,接收端只执行 args[0] 那个裸的 "RESTORE"), | ||
| // 客户端却收到 +OK。现在整条命令按复制键流那种 RESP 数组发,并且 | ||
| // 要求回复。 | ||
| // 2. 就算数据送到了,也不能凭空删源键:目标可能回 -BUSYKEY(键已存在且 | ||
| // 没带 REPLACE)。"目标没收下、源已经删了"是数据丢失,比留下重复键更糟。 | ||
| // 所以只有确认 +OK 之后才删。timeout 是这次往返的上限,到点报错、源键不动。 | ||
| std::string target_reply; | ||
| if (!conn->send_command_and_wait(target_name, restore_args, timeout, target_reply)) { | ||
| return RespEncoder::encode_error( | ||
| "ERR MIGRATE timed out or could not reach the target node; the source key was kept"); | ||
| } | ||
| if (target_reply.compare(0, 3, "+OK") != 0) { | ||
| // 把对端的错误原样转给客户端(-BUSYKEY ... 之类),但别顺手删源键 | ||
| if (!target_reply.empty() && target_reply[0] == '-') { | ||
| return RespEncoder::encode_error("ERR target node replied: " + | ||
| target_reply.substr(1, target_reply.find("\r\n") - 1)); | ||
| } | ||
| return RespEncoder::encode_error("ERR target node returned an unexpected reply"); | ||
| } | ||
|
|
||
| // 目标确认收下了,才动源键(Redis 的 MIGRATE 语义:默认就是移动,不是复制) | ||
| const bool removed = GlobalStorage::instance().del(key); | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
The source value is read before the network wait, but this unconditional delete runs afterward. If another client updates or recreates the key while MIGRATE waits, the target receives the old value and this line deletes the newer source value. Prompt to fix with AI
|
||
| LOG_INFO(CLUSTER, "MIGRATE completed: key=%s -> %s:%d source_removed=%d", | ||
| key.c_str(), host.c_str(), port, removed ? 1 : 0); | ||
| return RespEncoder::encode_simple_string("OK"); | ||
| } | ||
|
|
||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -398,6 +398,34 @@ async def test_replication(harness: ClusterTestHarness, r: TestResults): | |
| get_from_a == "from_master_a", f"got: {get_from_a}") | ||
|
|
||
|
|
||
| async def test_migrate(harness: ClusterTestHarness, r: TestResults): | ||
| """Test 4b: CLUSTER MIGRATE 真的把键搬走(目标收得到、源不再留副本).""" | ||
| print("\n── Test 4b: MIGRATE ──") | ||
|
|
||
| cli_a = harness.servers[16379].client | ||
| cli_b = harness.servers[16380].client | ||
|
|
||
| await cli_a.execute("SET", "mg_key", "mg_value") | ||
|
|
||
| # Redis 语法:MIGRATE host port key destination-db timeout [REPLACE] | ||
| mg = await cli_a.execute("MIGRATE", "127.0.0.1", "16380", "mg_key", "0", "3000", "REPLACE") | ||
| r.record("CLUSTER MIGRATE A→B 回 OK", mg == "OK", f"got: {mg}") | ||
|
Comment on lines
+411
to
+412
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. |
||
|
|
||
| await asyncio.sleep(0.5) | ||
|
|
||
| # 这两条是这条命令的真实契约。改动前它俩都不成立:RESTORE 的参数被拆成多个 | ||
| # 总线参数、对端只执行裸的 "RESTORE",所以键根本没搬过去,而源节点也从不删。 | ||
| gone_from_a = await cli_a.execute("GET", "mg_key") | ||
| r.record("源节点上这个键已经不在了", | ||
| gone_from_a is None or gone_from_a == "", | ||
| f"source still returns: {gone_from_a}") | ||
|
|
||
| arrived_on_b = await cli_b.execute("GET", "mg_key") | ||
| r.record("目标节点读得到搬过去的值", | ||
| arrived_on_b == "mg_value", | ||
| f"target returns: {arrived_on_b}") | ||
|
|
||
|
|
||
| async def test_failover_detection(harness: ClusterTestHarness, r: TestResults): | ||
| """Test 5: Node failure detection via CLUSTER FAIL.""" | ||
| print("\n── Test 5: Failover Detection ──") | ||
|
|
@@ -467,6 +495,7 @@ async def main(): | |
| await test_slot_assignment(harness, results) | ||
| await test_data_operations(harness, results) | ||
| await test_replication(harness, results) | ||
| await test_migrate(harness, results) | ||
| await test_failover_detection(harness, results) | ||
| await test_graceful_shutdown(harness, results) | ||
| except Exception as e: | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
This claims serialization uses byte-count-prefixed records and preserves embedded newlines, but the implementation emits newline-delimited values and RESTORE parses them with getline, so values containing newlines are truncated or misparsed.
Prompt to fix with AI