Hard resets triggered by installing Flatpak app on microSD

This is clearly a continuation of this thread:

Since then I have got rid of the questionable practice of storing swap on microSD on Librem 5, which made life better. I have also upgraded to Crimson. But the problem is still there: high IO on microSD crashes the phone.

In particular, now I need to install a Flatpak app onto an microSD installation. To be more precise, I have a LUKS inside LVM on a microSD. I have tried different tricks to limit IO, but each time the phone reboots with a short loud vibration.

I also tried limiting IO bandwidth with cgroups as following, but the phone still crashes.

systemd-run --pty \
            -p 'IOWriteBandwidthMax=/dev/dm-2 2M' \
            -p 'IOReadBandwidthMax=/dev/dm-2 2M' \
            -p 'IOWriteBandwidthMax=/dev/dm-0 2M' \
            -p 'IOReadBandwidthMax=/dev/dm-0 2M' \
            -p 'IOWriteBandwidthMax=/dev/sda 2M' \
            -p 'IOReadBandwidthMax=/dev/sda 2M' \
            \
            -p 'IOWriteIOPSMax=/dev/dm-2 2K' \
            -p 'IOReadIOPSMax=/dev/dm-2 2K' \
            -p 'IOWriteIOPSMax=/dev/dm-0 2K' \
            -p 'IOReadIOPSMax=/dev/dm-0 2K' \
            -p 'IOWriteIOPSMax=/dev/sda 2K' \
            -p 'IOReadIOPSMax=/dev/sda 2K' \
            flatpak --installation=sdcard install "$APP_NAME"

I also tried to manually pause installation at intervals and resume with the following script only after dstat -tdD total,sda,mmcblk0 5 reported almost no IO activity, but I have not succeeded:

pkill -CONT -e -f /bin/flatpak
sleep 20
pkill -STOP -e -f /bin/flatpak

I have repeatedly tried to install app this way more than a dozen of times in a row with no success. This is frustrating.

UPD: Another symptom is that the phone’s screen often freezes when there is such high IO. When it freezes like that, it has a chance to reset with a buzz, as described above.

I also noticed, that when I write to the same microSD from a laptop using a USB adapter, there is a USB traffic (as reported by usbmon) for a long time after block device IO stops being reported. The block device can not be unmounted until the USB traffic stops. Sadly, I have not found a tool to similarly monitor bus IO on microSD reader in Librem 5.

For fault isolation purposes, is it possible to install the app on the exact same µSD card but without LUKS and LVM i.e. vanilla file system? Or keep the volume arrangements as they are but try a different µSD card?

Presumably just some host-side caching?

Yes, but this is strange. We have two caches between a userspace program and hardware:

  • Delayed block device writing. A “write” system call returns before actual write finishes. The data is cached by kernel and gradually written to a block device. This is best illustrated by the iotop program, which displays separately IO between userspace and kernel and between kernel and block device.
  • A strange bus cache. Kernel’s block device engine thinks that data is already written, but there is another cache somewhere between the block device abstraction and the bus.

I focus on these caching peculiarities because it seems like phone crashes as soon as one of these caches fills. I have just managed to finally install the app after God knows how many attempts by manually pausing the process and waiting for scheduled writes to finish. Sadly, this is not a solution, but an ugly workaround.

Good idea, I’ll try it once I find my spare microSD. The same problem happens when copying large amount of files with rsync without bandwidth limit (i.e. the --bwlimit argument), so it will be easy to reproduce.

I’ve been running my flatpaks (and waydroid) from a 1TB microSD and I’ve never had any issues. The microSD is encrypted with LUKS. I’ve ran rsync backups from it, plus the native pureos backup and also rsync copies between internal storage and the microSD both ways. I’ve done 100GB+ copies via wired ethernet to and from the microSD card into my nas as well.

I’d suggest checking io usage real time (maybe via serial) and, yes, as already suggested, trying another card.

I remember changing the microSD card at least once. Maybe, it depends on manufacturer, its declared IO speed, or something like that. :thinking:

The first thing to do when this happens is to check the files in /var/lib/systemd/pstore to see what exactly made the kernel unhappy.

Librem 5’s microSD reader is a USB device, so the exact same tools apply.

I see /var/lib/systemd/pstore/dmesg-ramoops-0 file with today’s date in modification timestamp which ends with this:

dmesg-ramoops-0
<3>[ 4108.814536] INFO: task phoc:1120 blocked for more than 120 seconds.
<3>[ 4108.821051]       Not tainted 6.12.0-1-librem5 #2
<3>[ 4108.825961] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
<6>[ 4108.834006] task:phoc            state:D stack:0     pid:1120  tgid:1120  ppid:1      flags:0x00000004
<6>[ 4108.843444] Call trace:
<6>[ 4108.846069]  __switch_to+0xf0/0x150
<6>[ 4108.849687]  __schedule+0x2f0/0xc60
<6>[ 4108.853380]  schedule+0x3c/0xd8
<6>[ 4108.856642]  io_schedule+0x44/0x68
<6>[ 4108.860106]  folio_wait_bit_common+0x174/0x3d0
<6>[ 4108.864649]  folio_wait_bit+0x20/0x38
<6>[ 4108.868414]  folio_wait_writeback+0x54/0xc8
<6>[ 4108.872673]  migrate_pages_batch+0x248/0x1080
<6>[ 4108.877095]  migrate_pages_sync+0x158/0x230
<6>[ 4108.881395]  migrate_pages+0xac0/0xb80
<6>[ 4108.885207]  __alloc_contig_migrate_range+0xfc/0x368
<6>[ 4108.890249]  alloc_contig_range_noprof+0x190/0x580
<6>[ 4108.895144]  __cma_alloc+0x140/0x640
<6>[ 4108.898794]  cma_alloc+0x28/0x40
<6>[ 4108.902084]  cma_alloc_aligned+0x48/0x78
<6>[ 4108.906111]  dma_alloc_contiguous+0x38/0x58
<6>[ 4108.910377]  __dma_direct_alloc_pages.constprop.0+0xb4/0x298
<6>[ 4108.916142]  dma_direct_alloc+0x74/0x3a0
<6>[ 4108.920126]  dma_alloc_attrs+0x94/0x240
<6>[ 4108.924039]  drm_gem_dma_create+0xb0/0x168
<6>[ 4108.928239]  drm_gem_dma_create_with_handle+0x30/0xd8
<6>[ 4108.933366]  drm_gem_dma_dumb_create+0x40/0x58
<6>[ 4108.937912]  drm_mode_create_dumb_ioctl+0x98/0xc0
<6>[ 4108.942696]  drm_ioctl_kernel+0xcc/0x148
<6>[ 4108.946684]  drm_ioctl+0x268/0x560
<6>[ 4108.950202]  __arm64_sys_ioctl+0xbc/0x100
<6>[ 4108.954276]  invoke_syscall+0x4c/0xf8
<6>[ 4108.958001]  el0_svc_common.constprop.0+0x48/0xf0
<6>[ 4108.962809]  do_el0_svc+0x24/0x38
<6>[ 4108.966186]  el0_svc+0x2c/0x80
<6>[ 4108.969292]  el0t_64_sync_handler+0xc8/0xe0
<6>[ 4108.973577]  el0t_64_sync+0x190/0x198
<0>[ 4108.977512] Kernel panic - not syncing: hung_task: blocked tasks
<4>[ 4108.983580] CPU: 2 UID: 0 PID: 45 Comm: khungtaskd Not tainted 6.12.0-1-librem5 #2
<4>[ 4108.991250] Hardware name: Purism Librem 5r4 (DT)
<4>[ 4108.996045] Call trace:
<4>[ 4108.998541]  dump_backtrace+0xa0/0x128
<4>[ 4109.002348]  show_stack+0x20/0x38
<4>[ 4109.005714]  dump_stack_lvl+0x38/0x90
<4>[ 4109.009467]  dump_stack+0x18/0x28
<4>[ 4109.012833]  panic+0x3b8/0x3d8
<4>[ 4109.015938]  watchdog+0x2c0/0x580
<4>[ 4109.019347]  kthread+0x11c/0x128
<4>[ 4109.022626]  ret_from_fork+0x10/0x20
<2>[ 4109.026252] SMP: stopping secondary CPUs
<0>[ 4109.030771] Kernel Offset: disabled
<0>[ 4109.034308] CPU features: 0x00,00000080,00200000,4200421b
<0>[ 4109.039755] Memory Limit: none

ECC: 36 Corrected bytes, 0 unrecoverable blocks

We configure the kernel to force a panic when there are hung tasks detected, as in my experience these usually signify a problem that’s unlikely to resolve on its own and leaves the device in a state where it’s unable to function properly. Looks like in your case SD I/O is so slow that combined with the aggressive caching behavior your seem to be seeing, perhaps due to all the layering on top of the block device, it causes some very long delays with tasks waiting for file cache pages to be migrated onto the card before being able to allocate themselves, so the hung task detector kicks in. You can disable panics on hung tasks with sysctl.

That said, you are clearly running an out-of-date system. Recent wlroots package wouldn’t make phoc reallocate its swapchains on display unblank anymore, so this particular case wouldn’t trigger a CMA allocation at all, which is what hung for two minutes there in this trace you pasted. You’re also running a 6.12 kernel while we’re already on 6.18, and I know that there’s been a lot of work on folio that recently happened upstream, so I wouldn’t be surprised if this was already improved/fixed. I believe some of the kernel config changes we did recently could also help there.

Yea, listen to dos and keep your phone updated.

This is a promise Purism has been keeping consistently. The phone is way better now than 3 years ago, when I started daily driving it. =)

The last time I updated was indeed several months ago :thinking:

The kernel is indeed old, but I don’t see updates for wlroots. The currently installed version is 0.16.2-2pureos4:

Desired=Unknown/Install/Remove/Purge/Hold
| Status=Not/Inst/Conf-files/Unpacked/halF-conf/Half-inst/trig-aWait/Trig-pend
|/ Err?=(none)/Reinst-required (Status,Err: uppercase=bad)
||/ Name               Version         Architecture Description
+++-==================-===============-============-===================================================
ii  libwlroots11:arm64 0.16.2-2pureos4 arm64        Modular wayland compositor library - shared library

I probably need to buy a faster microSD. But I have read on this forum that read speed of the card reader itself in Librem 5 is limited. :thinking:

By the way, the microSD is reported by udevadm as following:

udevadm info -a -n /dev/sda
  looking at parent device '/devices/platform/soc@0/38200000.usb/xhci-hcd.4.auto/usb1/1-1/1-1.1/1-1.1:1.0/host0/target0:0:0/0:0:0:0':
    KERNELS=="0:0:0:0"
    SUBSYSTEMS=="scsi"
    DRIVERS=="sd"
    ATTRS{blacklist}=="INQUIRY_36 IGN_MEDIA_CHANGE"
    ATTRS{cdl_enable}=="0"
    ATTRS{cdl_supported}=="0"
    ATTRS{delete}=="(not readable)"
    ATTRS{device_blocked}=="0"
    ATTRS{device_busy}=="0"
    ATTRS{eh_timeout}=="10"
    ATTRS{evt_capacity_change_reported}=="0"
    ATTRS{evt_inquiry_change_reported}=="0"
    ATTRS{evt_lun_change_reported}=="0"
    ATTRS{evt_media_change}=="0"
    ATTRS{evt_mode_parameter_change_reported}=="0"
    ATTRS{evt_soft_threshold_reached}=="0"
    ATTRS{inquiry}==""
    ATTRS{iocounterbits}=="32"
    ATTRS{iodone_cnt}=="0x54c47"
    ATTRS{ioerr_cnt}=="0x4"
    ATTRS{iorequest_cnt}=="0x54c47"
    ATTRS{iotmo_cnt}=="0x0"
    ATTRS{max_sectors}=="240"
    ATTRS{model}=="Ultra HS-SD/MMC "
    ATTRS{power/autosuspend_delay_ms}=="500"
    ATTRS{power/control}=="auto"
    ATTRS{power/runtime_active_time}=="9641290"
    ATTRS{power/runtime_status}=="suspended"
    ATTRS{power/runtime_suspended_time}=="15743937"
    ATTRS{queue_depth}=="1"
    ATTRS{queue_type}=="none"
    ATTRS{rescan}=="(not readable)"
    ATTRS{rev}=="2.09"
    ATTRS{scsi_level}=="0"
    ATTRS{state}=="running"
    ATTRS{timeout}=="30"
    ATTRS{type}=="0"
    ATTRS{vendor}=="Generic "

In future, please don’t report issues that happen on severely outdated software - or at least not without clearly mentioning it. It just wastes everyone’s time, regardless of whether the reported things have been already fixed or not.

This is pretty solid advice on any device regardless of operating system, nature of device, …

Maybe the use of LVM is a bit niche in this context though. How many Librem 5 users would be using LVM on the µSD card?

I haven’t personally experienced this problem but my bulk updates to the content on the µSD card tend to be via the local network (sshfs), so perhaps too slow to get enough backlog to see this problem or maybe just not enough total MB being moved at once. In my case neither LVM nor LUKS is in use (not bothering with LUKS because 99.9% of the content is public content anyway).

Is it possible to lengthen the timeout? (while still keeping the panic if the timeout is reached)