Skip to content

Commit 0a4c604

Browse files
bingjingcgregkh
authored andcommitted
md: fix a potential deadlock of raid5/raid10 reshape
[ Upstream commit 8876391 ] There is a potential deadlock if mount/umount happens when raid5_finish_reshape() tries to grow the size of emulated disk. How the deadlock happens? 1) The raid5 resync thread finished reshape (expanding array). 2) The mount or umount thread holds VFS sb->s_umount lock and tries to write through critical data into raid5 emulated block device. So it waits for raid5 kernel thread handling stripes in order to finish it I/Os. 3) In the routine of raid5 kernel thread, md_check_recovery() will be called first in order to reap the raid5 resync thread. That is, raid5_finish_reshape() will be called. In this function, it will try to update conf and call VFS revalidate_disk() to grow the raid5 emulated block device. It will try to acquire VFS sb->s_umount lock. The raid5 kernel thread cannot continue, so no one can handle mount/ umount I/Os (stripes). Once the write-through I/Os cannot be finished, mount/umount will not release sb->s_umount lock. The deadlock happens. The raid5 kernel thread is an emulated block device. It is responible to handle I/Os (stripes) from upper layers. The emulated block device should not request any I/Os on itself. That is, it should not call VFS layer functions. (If it did, it will try to acquire VFS locks to guarantee the I/Os sequence.) So we have the resync thread to send resync I/O requests and to wait for the results. For solving this potential deadlock, we can put the size growth of the emulated block device as the final step of reshape thread. 2017/12/29: Thanks to Guoqing Jiang <gqjiang@suse.com>, we confirmed that there is the same deadlock issue in raid10. It's reproducible and can be fixed by this patch. For raid10.c, we can remove the similar code to prevent deadlock as well since they has been called before. Reported-by: Alex Wu <alexwu@synology.com> Reviewed-by: Alex Wu <alexwu@synology.com> Reviewed-by: Chung-Chiang Cheng <cccheng@synology.com> Signed-off-by: BingJing Chang <bingjingc@synology.com> Signed-off-by: Shaohua Li <sh.li@alibaba-inc.com> Signed-off-by: Sasha Levin <alexander.levin@microsoft.com> Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
1 parent 2565b27 commit 0a4c604

3 files changed

Lines changed: 15 additions & 14 deletions

File tree

‎drivers/md/md.c‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8522,6 +8522,19 @@ void md_do_sync(struct md_thread *thread)
85228522
set_mask_bits(&mddev->sb_flags, 0,
85238523
BIT(MD_SB_CHANGE_PENDING) | BIT(MD_SB_CHANGE_DEVS));
85248524

8525+
if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
8526+
!test_bit(MD_RECOVERY_INTR, &mddev->recovery) &&
8527+
mddev->delta_disks > 0 &&
8528+
mddev->pers->finish_reshape &&
8529+
mddev->pers->size &&
8530+
mddev->queue) {
8531+
mddev_lock_nointr(mddev);
8532+
md_set_array_sectors(mddev, mddev->pers->size(mddev, 0, 0));
8533+
mddev_unlock(mddev);
8534+
set_capacity(mddev->gendisk, mddev->array_sectors);
8535+
revalidate_disk(mddev->gendisk);
8536+
}
8537+
85258538
spin_lock(&mddev->lock);
85268539
if (!test_bit(MD_RECOVERY_INTR, &mddev->recovery)) {
85278540
/* We completed so min/max setting can be forgotten if used. */

‎drivers/md/raid10.c‎

Lines changed: 1 addition & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -4693,17 +4693,11 @@ static void raid10_finish_reshape(struct mddev *mddev)
46934693
return;
46944694

46954695
if (mddev->delta_disks > 0) {
4696-
sector_t size = raid10_size(mddev, 0, 0);
4697-
md_set_array_sectors(mddev, size);
46984696
if (mddev->recovery_cp > mddev->resync_max_sectors) {
46994697
mddev->recovery_cp = mddev->resync_max_sectors;
47004698
set_bit(MD_RECOVERY_NEEDED, &mddev->recovery);
47014699
}
4702-
mddev->resync_max_sectors = size;
4703-
if (mddev->queue) {
4704-
set_capacity(mddev->gendisk, mddev->array_sectors);
4705-
revalidate_disk(mddev->gendisk);
4706-
}
4700+
mddev->resync_max_sectors = mddev->array_sectors;
47074701
} else {
47084702
int d;
47094703
rcu_read_lock();

‎drivers/md/raid5.c‎

Lines changed: 1 addition & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -8001,13 +8001,7 @@ static void raid5_finish_reshape(struct mddev *mddev)
80018001

80028002
if (!test_bit(MD_RECOVERY_INTR, &mddev->recovery)) {
80038003

8004-
if (mddev->delta_disks > 0) {
8005-
md_set_array_sectors(mddev, raid5_size(mddev, 0, 0));
8006-
if (mddev->queue) {
8007-
set_capacity(mddev->gendisk, mddev->array_sectors);
8008-
revalidate_disk(mddev->gendisk);
8009-
}
8010-
} else {
8004+
if (mddev->delta_disks <= 0) {
80118005
int d;
80128006
spin_lock_irq(&conf->device_lock);
80138007
mddev->degraded = raid5_calc_degraded(conf);

0 commit comments

Comments
 (0)