<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic lx2160ardb: kernel reports problems during heavy disk I/O with PCIe/SATA adapters in Layerscape</title>
    <link>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/931124#M4474</link>
    <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;P&gt;Hello,&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Currently I'm trying to do some benchmarking with the LX2160ARDB and SATA SSDs. For this I have two PCIe-4x-to-4-port-SATA3.0 cards with Marvel 88SE9230 RAID controllers, that are plugged into the LX2160ARDB's PCIe slots (J23 and J28). Each has 4 Samsung 860 EVO SSDs connected for a total of 8 SSDs. The testing is done with Linux built from Yocto, btrfs raid0 and fio.&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Sometimes while writing data to the disks with fio I'm seeing various kernel errors, and noticable delays/hangs which interrupt the writing process for multiple seconds or even minutes. Some examples:&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;BLOCKQUOTE class="jive_macro_quote jive-quote jive_text_macro"&gt;&lt;P&gt;root@lx2160ardb:~# fio --name=x --directory=/mnt/btrfspool --numjobs=1 --size=3T --filesize=1G --nrfiles=3K --openfiles=1 --file_service_type=sequential --rw=write --ioengine=sync --bs=1M --direct=0 --fallocate=posix --zero_buffers=0 --write_bw_log=x&lt;BR /&gt;x: (g=0): rw=write, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=sync, iodepth=1&lt;BR /&gt;fio-3.12-dirty&lt;BR /&gt;Starting 1 process&lt;BR /&gt;x: Laying out IO files (3072 files / total 3145728MiB)&lt;BR /&gt;Jobs: 1 (f=1): [W(1)][0.6%][&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 1204.410448] mmc1: Timeout waiting for hardware cmd interrupt.&lt;BR /&gt;[ 1204.416186] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========&lt;BR /&gt;[ 1204.422615] mmc1: sdhci: Sys addr:&amp;nbsp; 0x5cc80000 | Version:&amp;nbsp; 0x00002202&lt;BR /&gt;[ 1204.429043] mmc1: sdhci: Blk size:&amp;nbsp; 0x00000200 | Blk cnt:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 1204.435471] mmc1: sdhci: Argument:&amp;nbsp; 0x00010000 | Trn mode: 0x00000033&lt;BR /&gt;[ 1204.441899] mmc1: sdhci: Present:&amp;nbsp;&amp;nbsp; 0x01f80008 | Host ctl: 0x0000003c&lt;BR /&gt;[ 1204.448327] mmc1: sdhci: Power:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000007 | Blk gap:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 1204.454755] mmc1: sdhci: Wake-up:&amp;nbsp;&amp;nbsp; 0x00000000 | Clock:&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000208&lt;BR /&gt;[ 1204.461183] mmc1: sdhci: Timeout:&amp;nbsp;&amp;nbsp; 0x0000000e | Int stat: 0x00000001&lt;BR /&gt;[ 1204.467610] mmc1: sdhci: Int enab:&amp;nbsp; 0x037f100f | Sig enab: 0x037f100b&lt;BR /&gt;[ 1204.474038] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00002202&lt;BR /&gt;[ 1204.480465] mmc1: sdhci: Caps:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x34fa0000 | Caps_1:&amp;nbsp;&amp;nbsp; 0x0000af00&lt;BR /&gt;[ 1204.486893] mmc1: sdhci: Cmd:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000d1a | Max curr: 0x00000000&lt;BR /&gt;[ 1204.493321] mmc1: sdhci: Resp[0]:&amp;nbsp;&amp;nbsp; 0x00000900 | Resp[1]:&amp;nbsp; 0xffffff8d&lt;BR /&gt;[ 1204.499748] mmc1: sdhci: Resp[2]:&amp;nbsp;&amp;nbsp; 0x320f5903 | Resp[3]:&amp;nbsp; 0x00000900&lt;BR /&gt;[ 1204.506175] mmc1: sdhci: Host ctl2: 0x00000080&lt;BR /&gt;[ 1204.510607] mmc1: sdhci: ADMA Err:&amp;nbsp; 0x00000000 | ADMA Ptr: 0x00000000f9c9820c&lt;BR /&gt;[ 1204.517729] mmc1: sdhci: ============================================&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 3565.222434] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:&amp;nbsp; &amp;nbsp;&lt;BR /&gt;[ 3565.228521] rcu: &amp;nbsp;&amp;nbsp; &amp;nbsp;(detected by 14, t=5252 jiffies, g=610161, q=2482)&lt;BR /&gt;[ 3565.234867] rcu: All QSes seen, last rcu_preempt kthread activity 5250 (4295783547-4295778297), jiffies_till_next_fqs=1, root -&amp;gt;qsmask 0x0&lt;BR /&gt;[ 3565.247284] kworker/u32:27&amp;nbsp; R&amp;nbsp; running task&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp; 3372&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 2 0x0000002a&lt;BR /&gt;[ 3565.254333] Workqueue: btrfs-endio-write btrfs_endio_write_helper&lt;BR /&gt;[ 3565.260414] Call trace:&lt;BR /&gt;[ 3565.262851]&amp;nbsp; dump_backtrace+0x0/0x158&lt;BR /&gt;[ 3565.266502]&amp;nbsp; show_stack+0x14/0x20&lt;BR /&gt;[ 3565.269805]&amp;nbsp; sched_show_task+0x13c/0x168&lt;BR /&gt;[ 3565.273716]&amp;nbsp; rcu_check_callbacks+0x7e0/0x850&lt;BR /&gt;[ 3565.277975]&amp;nbsp; update_process_times+0x2c/0x70&lt;BR /&gt;[ 3565.282147]&amp;nbsp; tick_sched_handle.isra.5+0x3c/0x50&lt;BR /&gt;[ 3565.286664]&amp;nbsp; tick_sched_timer+0x48/0x98&lt;BR /&gt;[ 3565.290487]&amp;nbsp; __hrtimer_run_queues+0x118/0x1a8&lt;BR /&gt;[ 3565.294831]&amp;nbsp; hrtimer_interrupt+0xe4/0x240&lt;BR /&gt;[ 3565.298829]&amp;nbsp; arch_timer_handler_phys+0x2c/0x38&lt;BR /&gt;[ 3565.303262]&amp;nbsp; handle_percpu_devid_irq+0x80/0x138&lt;BR /&gt;[ 3565.307779]&amp;nbsp; generic_handle_irq+0x24/0x38&lt;BR /&gt;[ 3565.311776]&amp;nbsp; __handle_domain_irq+0x60/0xb8&lt;BR /&gt;[ 3565.315860]&amp;nbsp; gic_handle_irq+0x7c/0x178&lt;BR /&gt;[ 3565.319596]&amp;nbsp; el1_irq+0xb0/0x128&lt;BR /&gt;[ 3565.322726]&amp;nbsp; queued_spin_lock_slowpath+0x230/0x2a8&lt;BR /&gt;[ 3565.327505]&amp;nbsp; queued_read_lock_slowpath+0x118/0x120&lt;BR /&gt;[ 3565.332285]&amp;nbsp; _raw_read_lock+0x44/0x48&lt;BR /&gt;[ 3565.335935]&amp;nbsp; btrfs_tree_read_lock+0x40/0x130&lt;BR /&gt;[ 3565.340193]&amp;nbsp; btrfs_search_slot+0x6d4/0x8a0&lt;BR /&gt;[ 3565.344278]&amp;nbsp; btrfs_lookup_csum+0x5c/0x188&lt;BR /&gt;[ 3565.348275]&amp;nbsp; btrfs_csum_file_blocks+0x214/0x580&lt;BR /&gt;[ 3565.352794]&amp;nbsp; add_pending_csums+0x64/0x98&lt;BR /&gt;[ 3565.356705]&amp;nbsp; btrfs_finish_ordered_io+0x2c0/0x810&lt;BR /&gt;[ 3565.361309]&amp;nbsp; finish_ordered_fn+0x10/0x18&lt;BR /&gt;[ 3565.365219]&amp;nbsp; normal_work_helper+0x228/0x240&lt;BR /&gt;[ 3565.369390]&amp;nbsp; btrfs_endio_write_helper+0x10/0x18&lt;BR /&gt;[ 3565.373909]&amp;nbsp; process_one_work+0x1e0/0x318&lt;BR /&gt;[ 3565.377907]&amp;nbsp; worker_thread+0x40/0x428&lt;BR /&gt;[ 3565.381557]&amp;nbsp; kthread+0x124/0x128&lt;BR /&gt;[ 3565.384772]&amp;nbsp; ret_from_fork+0x10/0x18&lt;BR /&gt;[ 3565.388337] rcu: rcu_preempt kthread starved for 5250 jiffies! g610161 f0x2 RCU_GP_WAIT_FQS(5) -&amp;gt;state=0x200 -&amp;gt;cpu=7&lt;BR /&gt;[ 3565.398843] rcu: RCU grace-period kthread stack dump:&lt;BR /&gt;[ 3565.403881] rcu_preempt&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; R&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp; 10&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 2 0x00000028&lt;BR /&gt;[ 3565.409355] Call trace:&lt;BR /&gt;[ 3565.411789]&amp;nbsp; __switch_to+0xa0/0xe0&lt;BR /&gt;[ 3565.415179]&amp;nbsp; __schedule+0x1e0/0x5c0&lt;BR /&gt;[ 3565.418655]&amp;nbsp; schedule+0x38/0xa0&lt;BR /&gt;[ 3565.421784]&amp;nbsp; schedule_timeout+0x198/0x338&lt;BR /&gt;[ 3565.425783]&amp;nbsp; rcu_gp_kthread+0x428/0x800&lt;BR /&gt;[ 3565.429607]&amp;nbsp; kthread+0x124/0x128&lt;BR /&gt;[ 3565.432822]&amp;nbsp; ret_from_fork+0x10/0x18&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 4510.522428] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:&lt;BR /&gt;[ 4510.528509] rcu: &amp;nbsp;&amp;nbsp; &amp;nbsp;(detected by 15, t=241578 jiffies, g=610161, q=96607)&lt;BR /&gt;[ 4510.535111] rcu: All QSes seen, last rcu_preempt kthread activity 241578 (4296019875-4295778297), jiffies_till_next_fqs=1, root -&amp;gt;qsmask 0x0&lt;BR /&gt;[ 4510.547701] swapper/15&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; R&amp;nbsp; running task&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 1 0x00000028&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;an example from another run:&lt;/P&gt;&lt;BLOCKQUOTE class="jive_macro_quote jive-quote jive_text_macro"&gt;&lt;P&gt;[ 4261.539921] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========&lt;BR /&gt;[ 4261.546349] mmc1: sdhci: Sys addr:&amp;nbsp; 0x5c580000 | Version:&amp;nbsp; 0x00002202&lt;BR /&gt;[ 4261.552777] mmc1: sdhci: Blk size:&amp;nbsp; 0x00000200 | Blk cnt:&amp;nbsp; 0x00000008&lt;BR /&gt;[ 4261.559204] mmc1: sdhci: Argument:&amp;nbsp; 0x00010000 | Trn mode: 0x00000033&lt;BR /&gt;[ 4261.565631] mmc1: sdhci: Present:&amp;nbsp;&amp;nbsp; 0x01f80008 | Host ctl: 0x0000003c&lt;BR /&gt;[ 4261.572059] mmc1: sdhci: Power:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000007 | Blk gap:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 4261.578487] mmc1: sdhci: Wake-up:&amp;nbsp;&amp;nbsp; 0x00000000 | Clock:&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000208&lt;BR /&gt;[ 4261.584914] mmc1: sdhci: Timeout:&amp;nbsp;&amp;nbsp; 0x0000000e | Int stat: 0x00000001&lt;BR /&gt;[ 4261.591342] mmc1: sdhci: Int enab:&amp;nbsp; 0x037f100f | Sig enab: 0x037f100b&lt;BR /&gt;[ 4261.597770] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00002202&lt;BR /&gt;[ 4261.604197] mmc1: sdhci: Caps:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x34fa0000 | Caps_1:&amp;nbsp;&amp;nbsp; 0x0000af00&lt;BR /&gt;[ 4261.610625] mmc1: sdhci: Cmd:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000d1a | Max curr: 0x00000000&lt;BR /&gt;[ 4261.617052] mmc1: sdhci: Resp[0]:&amp;nbsp;&amp;nbsp; 0x00000900 | Resp[1]:&amp;nbsp; 0x0000000d&lt;BR /&gt;[ 4261.623480] mmc1: sdhci: Resp[2]:&amp;nbsp;&amp;nbsp; 0x00000000 | Resp[3]:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 4261.629907] mmc1: sdhci: Host ctl2: 0x00000080&lt;BR /&gt;[ 4261.634339] mmc1: sdhci: ADMA Err:&amp;nbsp; 0x00000000 | ADMA Ptr: 0x00000000f9c9820c&lt;BR /&gt;[ 4261.641459] mmc1: sdhci: ============================================&lt;BR /&gt;[ 4261.648429] mmc1: card 0001 removed&lt;BR /&gt;[ 4271.982233] ata11.00: exception Emask 0x0 SAct 0xffc00bff SErr 0x0 action 0x6 frozen&lt;BR /&gt;[ 4271.989975] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4271.995330] ata11.00: cmd 61/80:00:80:4a:9c/00:00:2f:00:00/40 tag 0 ncq dma 65536 out&lt;BR /&gt;[ 4271.995330]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4271.996397] sd 3:0:0:0: [sdd] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.010534] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010537] ata11.00: failed command: WRITE FPDMA QUEUED&lt;/P&gt;&lt;P&gt;[ 4272.010542] ata11.00: cmd 61/00:08:00:45:9c/02:00:2f:00:00/40 tag 1 ncq dma 262144 out&lt;BR /&gt;[ 4272.010542]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;/P&gt;&lt;P&gt;[ 4272.010544] ata11.00: status: { DRDY }&lt;/P&gt;&lt;P&gt;fio: io_u error on file /mnt/btrfspool/x.0.3042: Input/output error: write offset=195035136, buflen=1048576&lt;BR /&gt;[ 4272.010546] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010550] ata11.00: cmd 61/00:10:00:48:9c/01:00:2f:00:00/40 tag 2 ncq dma 131072 out&lt;BR /&gt;[ 4272.010550]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010552] ata11.00: status: { DRDY }&lt;/P&gt;&lt;P&gt;[ 4272.010554] ata11.00: failed command: WRITE FPDMA QUEUED&lt;/P&gt;&lt;P&gt;[ 4272.010558] ata11.00: cmd 61/00:18:00:4b:9c/01:00:2f:00:00/40 tag 3 ncq dma 131072 out&lt;/P&gt;&lt;P&gt;fio: io_u error on file /mnt/btrfspool/x.0.3042: Input/output error: write offset=240123904, buflen=1048576&lt;BR /&gt;[ 4272.010558]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010559] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010561] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010565] ata11.00: cmd 61/80:20:80:44:9c/00:00:2f:00:00/40 tag 4 ncq dma 65536 out&lt;BR /&gt;[ 4272.010565]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010566] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010568] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010573] ata11.00: cmd 61/00:28:00:4c:9c/01:00:2f:00:00/40 tag 5 ncq dma 131072 out&lt;BR /&gt;[ 4272.010573]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010574] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010579] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.011649] sd 2:0:0:0: [sdc] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011654] sd 2:0:0:0: [sdc] tag#31 CDB: opcode=0x2a 2a 00 2f 9b e8 80 00 00 80 00&lt;BR /&gt;[ 4272.011657] print_req_error: I/O error, dev sdc, sector 798746752&lt;BR /&gt;[ 4272.011659] sd 1:0:0:0: [sdb] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011664] BTRFS error (device sda): bdev /dev/sdc errs: wr 1, rd 0, flush 0, corrupt 0, gen 0&lt;BR /&gt;[ 4272.011665] sd 1:0:0:0: [sdb] tag#31 CDB: opcode=0x2a 2a 00 2f 9b e9 80 00 00 80 00&lt;BR /&gt;[ 4272.011669] BTRFS warning (device sda): direct IO failed ino 4323 rw 1,34817 sector 0x2f9be900 len 0 err no 10&lt;BR /&gt;[ 4272.011671] print_req_error: I/O error, dev sdb, sector 798747008&lt;BR /&gt;[ 4272.011676] BTRFS error (device sda): bdev /dev/sdb errs: wr 1, rd 0, flush 0, corrupt 0, gen 0&lt;BR /&gt;[ 4272.011680] BTRFS warning (device sda): direct IO failed ino 4323 rw 1,34817 sector 0x2f9bea00 len 0 err no 10&lt;BR /&gt;[ 4272.011705] sd 2:0:0:0: [sdc] tag#30 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011708] sd 2:0:0:0: [sdc] tag#30 CDB: opcode=0x2a 2a 00 2f 9b d2 00 00 0a 00 00&lt;BR /&gt;[ 4272.011710] print_req_error: I/O error, dev sdc, sector 798740992&lt;BR /&gt;[ 4272.011713] BTRFS error (device sda): bdev /dev/sdc errs: wr 2, rd 0, flush 0, corrupt 0, gen 0&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Some additional information that might be useful:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;I started out with the Yocto state corresponding to LSDK-19.03 (as per yocto_2.7 tag in the qoriq-components/yocto-sdk repo). However this included the linux-qoriq kernel 4.19.26+gc0c214110624 which crashed at boot in ahci-qoriq driver's fsl_sata_errata_379364() function. That was fixed by upgrading to kernel 4.19.46+g1a4cab2c597d from LSDK-19.06. (some other components such as u-boot and management-complex firmware had to be upgraded to LSDK-19.06 state too to make the new kernel work)&lt;/LI&gt;&lt;LI&gt;I did not see such problems during disk I/O with only 4 SSDs connected directly to the board's SATA connectors. Also no problems with 4 SSDs connected to only 1 PCIe card, in either slot, as opposed to using 2 PCIe cards.&lt;/LI&gt;&lt;LI&gt;The same problems appeared when using 2 PCIe cards with 2 SSDs each (total 4 SSDs instead of 8).&lt;/LI&gt;&lt;LI&gt;We tried powering half of the SSDs with an external power supply, in case the internal one was too weak for 8 SSDs, but it doesn't seem to make a difference - errors appeared either way.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Does anyone have an idea what the problem could be? I'm pretty much lost at this point. It would be really nice to get the 8 SSDs to work reliably, but so far, no luck. Probably the way forward is to do more tests to rule out hardware problems with the PCIe cards or SSDs themselves.&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;But I don't understand why doing disk I/O would cause mmc1 timeouts. Yes, I'm booting from SD card, but that's mmc0 - mmc1 is (as far as I know) the internal eMMC which isn't being used at all (when mounting the btrfs raid0, only /dev/sd[a-h] are used).&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Thanks in advance for any hints.&lt;/P&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
    <pubDate>Wed, 17 Jul 2019 12:08:10 GMT</pubDate>
    <dc:creator>daniel_klauer</dc:creator>
    <dc:date>2019-07-17T12:08:10Z</dc:date>
    <item>
      <title>lx2160ardb: kernel reports problems during heavy disk I/O with PCIe/SATA adapters</title>
      <link>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/931124#M4474</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;P&gt;Hello,&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Currently I'm trying to do some benchmarking with the LX2160ARDB and SATA SSDs. For this I have two PCIe-4x-to-4-port-SATA3.0 cards with Marvel 88SE9230 RAID controllers, that are plugged into the LX2160ARDB's PCIe slots (J23 and J28). Each has 4 Samsung 860 EVO SSDs connected for a total of 8 SSDs. The testing is done with Linux built from Yocto, btrfs raid0 and fio.&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Sometimes while writing data to the disks with fio I'm seeing various kernel errors, and noticable delays/hangs which interrupt the writing process for multiple seconds or even minutes. Some examples:&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;BLOCKQUOTE class="jive_macro_quote jive-quote jive_text_macro"&gt;&lt;P&gt;root@lx2160ardb:~# fio --name=x --directory=/mnt/btrfspool --numjobs=1 --size=3T --filesize=1G --nrfiles=3K --openfiles=1 --file_service_type=sequential --rw=write --ioengine=sync --bs=1M --direct=0 --fallocate=posix --zero_buffers=0 --write_bw_log=x&lt;BR /&gt;x: (g=0): rw=write, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=sync, iodepth=1&lt;BR /&gt;fio-3.12-dirty&lt;BR /&gt;Starting 1 process&lt;BR /&gt;x: Laying out IO files (3072 files / total 3145728MiB)&lt;BR /&gt;Jobs: 1 (f=1): [W(1)][0.6%][&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 1204.410448] mmc1: Timeout waiting for hardware cmd interrupt.&lt;BR /&gt;[ 1204.416186] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========&lt;BR /&gt;[ 1204.422615] mmc1: sdhci: Sys addr:&amp;nbsp; 0x5cc80000 | Version:&amp;nbsp; 0x00002202&lt;BR /&gt;[ 1204.429043] mmc1: sdhci: Blk size:&amp;nbsp; 0x00000200 | Blk cnt:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 1204.435471] mmc1: sdhci: Argument:&amp;nbsp; 0x00010000 | Trn mode: 0x00000033&lt;BR /&gt;[ 1204.441899] mmc1: sdhci: Present:&amp;nbsp;&amp;nbsp; 0x01f80008 | Host ctl: 0x0000003c&lt;BR /&gt;[ 1204.448327] mmc1: sdhci: Power:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000007 | Blk gap:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 1204.454755] mmc1: sdhci: Wake-up:&amp;nbsp;&amp;nbsp; 0x00000000 | Clock:&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000208&lt;BR /&gt;[ 1204.461183] mmc1: sdhci: Timeout:&amp;nbsp;&amp;nbsp; 0x0000000e | Int stat: 0x00000001&lt;BR /&gt;[ 1204.467610] mmc1: sdhci: Int enab:&amp;nbsp; 0x037f100f | Sig enab: 0x037f100b&lt;BR /&gt;[ 1204.474038] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00002202&lt;BR /&gt;[ 1204.480465] mmc1: sdhci: Caps:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x34fa0000 | Caps_1:&amp;nbsp;&amp;nbsp; 0x0000af00&lt;BR /&gt;[ 1204.486893] mmc1: sdhci: Cmd:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000d1a | Max curr: 0x00000000&lt;BR /&gt;[ 1204.493321] mmc1: sdhci: Resp[0]:&amp;nbsp;&amp;nbsp; 0x00000900 | Resp[1]:&amp;nbsp; 0xffffff8d&lt;BR /&gt;[ 1204.499748] mmc1: sdhci: Resp[2]:&amp;nbsp;&amp;nbsp; 0x320f5903 | Resp[3]:&amp;nbsp; 0x00000900&lt;BR /&gt;[ 1204.506175] mmc1: sdhci: Host ctl2: 0x00000080&lt;BR /&gt;[ 1204.510607] mmc1: sdhci: ADMA Err:&amp;nbsp; 0x00000000 | ADMA Ptr: 0x00000000f9c9820c&lt;BR /&gt;[ 1204.517729] mmc1: sdhci: ============================================&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 3565.222434] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:&amp;nbsp; &amp;nbsp;&lt;BR /&gt;[ 3565.228521] rcu: &amp;nbsp;&amp;nbsp; &amp;nbsp;(detected by 14, t=5252 jiffies, g=610161, q=2482)&lt;BR /&gt;[ 3565.234867] rcu: All QSes seen, last rcu_preempt kthread activity 5250 (4295783547-4295778297), jiffies_till_next_fqs=1, root -&amp;gt;qsmask 0x0&lt;BR /&gt;[ 3565.247284] kworker/u32:27&amp;nbsp; R&amp;nbsp; running task&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp; 3372&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 2 0x0000002a&lt;BR /&gt;[ 3565.254333] Workqueue: btrfs-endio-write btrfs_endio_write_helper&lt;BR /&gt;[ 3565.260414] Call trace:&lt;BR /&gt;[ 3565.262851]&amp;nbsp; dump_backtrace+0x0/0x158&lt;BR /&gt;[ 3565.266502]&amp;nbsp; show_stack+0x14/0x20&lt;BR /&gt;[ 3565.269805]&amp;nbsp; sched_show_task+0x13c/0x168&lt;BR /&gt;[ 3565.273716]&amp;nbsp; rcu_check_callbacks+0x7e0/0x850&lt;BR /&gt;[ 3565.277975]&amp;nbsp; update_process_times+0x2c/0x70&lt;BR /&gt;[ 3565.282147]&amp;nbsp; tick_sched_handle.isra.5+0x3c/0x50&lt;BR /&gt;[ 3565.286664]&amp;nbsp; tick_sched_timer+0x48/0x98&lt;BR /&gt;[ 3565.290487]&amp;nbsp; __hrtimer_run_queues+0x118/0x1a8&lt;BR /&gt;[ 3565.294831]&amp;nbsp; hrtimer_interrupt+0xe4/0x240&lt;BR /&gt;[ 3565.298829]&amp;nbsp; arch_timer_handler_phys+0x2c/0x38&lt;BR /&gt;[ 3565.303262]&amp;nbsp; handle_percpu_devid_irq+0x80/0x138&lt;BR /&gt;[ 3565.307779]&amp;nbsp; generic_handle_irq+0x24/0x38&lt;BR /&gt;[ 3565.311776]&amp;nbsp; __handle_domain_irq+0x60/0xb8&lt;BR /&gt;[ 3565.315860]&amp;nbsp; gic_handle_irq+0x7c/0x178&lt;BR /&gt;[ 3565.319596]&amp;nbsp; el1_irq+0xb0/0x128&lt;BR /&gt;[ 3565.322726]&amp;nbsp; queued_spin_lock_slowpath+0x230/0x2a8&lt;BR /&gt;[ 3565.327505]&amp;nbsp; queued_read_lock_slowpath+0x118/0x120&lt;BR /&gt;[ 3565.332285]&amp;nbsp; _raw_read_lock+0x44/0x48&lt;BR /&gt;[ 3565.335935]&amp;nbsp; btrfs_tree_read_lock+0x40/0x130&lt;BR /&gt;[ 3565.340193]&amp;nbsp; btrfs_search_slot+0x6d4/0x8a0&lt;BR /&gt;[ 3565.344278]&amp;nbsp; btrfs_lookup_csum+0x5c/0x188&lt;BR /&gt;[ 3565.348275]&amp;nbsp; btrfs_csum_file_blocks+0x214/0x580&lt;BR /&gt;[ 3565.352794]&amp;nbsp; add_pending_csums+0x64/0x98&lt;BR /&gt;[ 3565.356705]&amp;nbsp; btrfs_finish_ordered_io+0x2c0/0x810&lt;BR /&gt;[ 3565.361309]&amp;nbsp; finish_ordered_fn+0x10/0x18&lt;BR /&gt;[ 3565.365219]&amp;nbsp; normal_work_helper+0x228/0x240&lt;BR /&gt;[ 3565.369390]&amp;nbsp; btrfs_endio_write_helper+0x10/0x18&lt;BR /&gt;[ 3565.373909]&amp;nbsp; process_one_work+0x1e0/0x318&lt;BR /&gt;[ 3565.377907]&amp;nbsp; worker_thread+0x40/0x428&lt;BR /&gt;[ 3565.381557]&amp;nbsp; kthread+0x124/0x128&lt;BR /&gt;[ 3565.384772]&amp;nbsp; ret_from_fork+0x10/0x18&lt;BR /&gt;[ 3565.388337] rcu: rcu_preempt kthread starved for 5250 jiffies! g610161 f0x2 RCU_GP_WAIT_FQS(5) -&amp;gt;state=0x200 -&amp;gt;cpu=7&lt;BR /&gt;[ 3565.398843] rcu: RCU grace-period kthread stack dump:&lt;BR /&gt;[ 3565.403881] rcu_preempt&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; R&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp; 10&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 2 0x00000028&lt;BR /&gt;[ 3565.409355] Call trace:&lt;BR /&gt;[ 3565.411789]&amp;nbsp; __switch_to+0xa0/0xe0&lt;BR /&gt;[ 3565.415179]&amp;nbsp; __schedule+0x1e0/0x5c0&lt;BR /&gt;[ 3565.418655]&amp;nbsp; schedule+0x38/0xa0&lt;BR /&gt;[ 3565.421784]&amp;nbsp; schedule_timeout+0x198/0x338&lt;BR /&gt;[ 3565.425783]&amp;nbsp; rcu_gp_kthread+0x428/0x800&lt;BR /&gt;[ 3565.429607]&amp;nbsp; kthread+0x124/0x128&lt;BR /&gt;[ 3565.432822]&amp;nbsp; ret_from_fork+0x10/0x18&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[ 4510.522428] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:&lt;BR /&gt;[ 4510.528509] rcu: &amp;nbsp;&amp;nbsp; &amp;nbsp;(detected by 15, t=241578 jiffies, g=610161, q=96607)&lt;BR /&gt;[ 4510.535111] rcu: All QSes seen, last rcu_preempt kthread activity 241578 (4296019875-4295778297), jiffies_till_next_fqs=1, root -&amp;gt;qsmask 0x0&lt;BR /&gt;[ 4510.547701] swapper/15&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; R&amp;nbsp; running task&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 1 0x00000028&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;[...]&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;an example from another run:&lt;/P&gt;&lt;BLOCKQUOTE class="jive_macro_quote jive-quote jive_text_macro"&gt;&lt;P&gt;[ 4261.539921] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========&lt;BR /&gt;[ 4261.546349] mmc1: sdhci: Sys addr:&amp;nbsp; 0x5c580000 | Version:&amp;nbsp; 0x00002202&lt;BR /&gt;[ 4261.552777] mmc1: sdhci: Blk size:&amp;nbsp; 0x00000200 | Blk cnt:&amp;nbsp; 0x00000008&lt;BR /&gt;[ 4261.559204] mmc1: sdhci: Argument:&amp;nbsp; 0x00010000 | Trn mode: 0x00000033&lt;BR /&gt;[ 4261.565631] mmc1: sdhci: Present:&amp;nbsp;&amp;nbsp; 0x01f80008 | Host ctl: 0x0000003c&lt;BR /&gt;[ 4261.572059] mmc1: sdhci: Power:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000007 | Blk gap:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 4261.578487] mmc1: sdhci: Wake-up:&amp;nbsp;&amp;nbsp; 0x00000000 | Clock:&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000208&lt;BR /&gt;[ 4261.584914] mmc1: sdhci: Timeout:&amp;nbsp;&amp;nbsp; 0x0000000e | Int stat: 0x00000001&lt;BR /&gt;[ 4261.591342] mmc1: sdhci: Int enab:&amp;nbsp; 0x037f100f | Sig enab: 0x037f100b&lt;BR /&gt;[ 4261.597770] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00002202&lt;BR /&gt;[ 4261.604197] mmc1: sdhci: Caps:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x34fa0000 | Caps_1:&amp;nbsp;&amp;nbsp; 0x0000af00&lt;BR /&gt;[ 4261.610625] mmc1: sdhci: Cmd:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 0x00000d1a | Max curr: 0x00000000&lt;BR /&gt;[ 4261.617052] mmc1: sdhci: Resp[0]:&amp;nbsp;&amp;nbsp; 0x00000900 | Resp[1]:&amp;nbsp; 0x0000000d&lt;BR /&gt;[ 4261.623480] mmc1: sdhci: Resp[2]:&amp;nbsp;&amp;nbsp; 0x00000000 | Resp[3]:&amp;nbsp; 0x00000000&lt;BR /&gt;[ 4261.629907] mmc1: sdhci: Host ctl2: 0x00000080&lt;BR /&gt;[ 4261.634339] mmc1: sdhci: ADMA Err:&amp;nbsp; 0x00000000 | ADMA Ptr: 0x00000000f9c9820c&lt;BR /&gt;[ 4261.641459] mmc1: sdhci: ============================================&lt;BR /&gt;[ 4261.648429] mmc1: card 0001 removed&lt;BR /&gt;[ 4271.982233] ata11.00: exception Emask 0x0 SAct 0xffc00bff SErr 0x0 action 0x6 frozen&lt;BR /&gt;[ 4271.989975] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4271.995330] ata11.00: cmd 61/80:00:80:4a:9c/00:00:2f:00:00/40 tag 0 ncq dma 65536 out&lt;BR /&gt;[ 4271.995330]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4271.996397] sd 3:0:0:0: [sdd] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.010534] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010537] ata11.00: failed command: WRITE FPDMA QUEUED&lt;/P&gt;&lt;P&gt;[ 4272.010542] ata11.00: cmd 61/00:08:00:45:9c/02:00:2f:00:00/40 tag 1 ncq dma 262144 out&lt;BR /&gt;[ 4272.010542]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;/P&gt;&lt;P&gt;[ 4272.010544] ata11.00: status: { DRDY }&lt;/P&gt;&lt;P&gt;fio: io_u error on file /mnt/btrfspool/x.0.3042: Input/output error: write offset=195035136, buflen=1048576&lt;BR /&gt;[ 4272.010546] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010550] ata11.00: cmd 61/00:10:00:48:9c/01:00:2f:00:00/40 tag 2 ncq dma 131072 out&lt;BR /&gt;[ 4272.010550]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010552] ata11.00: status: { DRDY }&lt;/P&gt;&lt;P&gt;[ 4272.010554] ata11.00: failed command: WRITE FPDMA QUEUED&lt;/P&gt;&lt;P&gt;[ 4272.010558] ata11.00: cmd 61/00:18:00:4b:9c/01:00:2f:00:00/40 tag 3 ncq dma 131072 out&lt;/P&gt;&lt;P&gt;fio: io_u error on file /mnt/btrfspool/x.0.3042: Input/output error: write offset=240123904, buflen=1048576&lt;BR /&gt;[ 4272.010558]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010559] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010561] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010565] ata11.00: cmd 61/80:20:80:44:9c/00:00:2f:00:00/40 tag 4 ncq dma 65536 out&lt;BR /&gt;[ 4272.010565]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010566] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010568] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.010573] ata11.00: cmd 61/00:28:00:4c:9c/01:00:2f:00:00/40 tag 5 ncq dma 131072 out&lt;BR /&gt;[ 4272.010573]&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; res 40/00:00:00:00:00/00:00:00:00:00/00 Emask 0x4 (timeout)&lt;BR /&gt;[ 4272.010574] ata11.00: status: { DRDY }&lt;BR /&gt;[ 4272.010579] ata11.00: failed command: WRITE FPDMA QUEUED&lt;BR /&gt;[ 4272.011649] sd 2:0:0:0: [sdc] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011654] sd 2:0:0:0: [sdc] tag#31 CDB: opcode=0x2a 2a 00 2f 9b e8 80 00 00 80 00&lt;BR /&gt;[ 4272.011657] print_req_error: I/O error, dev sdc, sector 798746752&lt;BR /&gt;[ 4272.011659] sd 1:0:0:0: [sdb] tag#31 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011664] BTRFS error (device sda): bdev /dev/sdc errs: wr 1, rd 0, flush 0, corrupt 0, gen 0&lt;BR /&gt;[ 4272.011665] sd 1:0:0:0: [sdb] tag#31 CDB: opcode=0x2a 2a 00 2f 9b e9 80 00 00 80 00&lt;BR /&gt;[ 4272.011669] BTRFS warning (device sda): direct IO failed ino 4323 rw 1,34817 sector 0x2f9be900 len 0 err no 10&lt;BR /&gt;[ 4272.011671] print_req_error: I/O error, dev sdb, sector 798747008&lt;BR /&gt;[ 4272.011676] BTRFS error (device sda): bdev /dev/sdb errs: wr 1, rd 0, flush 0, corrupt 0, gen 0&lt;BR /&gt;[ 4272.011680] BTRFS warning (device sda): direct IO failed ino 4323 rw 1,34817 sector 0x2f9bea00 len 0 err no 10&lt;BR /&gt;[ 4272.011705] sd 2:0:0:0: [sdc] tag#30 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;[ 4272.011708] sd 2:0:0:0: [sdc] tag#30 CDB: opcode=0x2a 2a 00 2f 9b d2 00 00 0a 00 00&lt;BR /&gt;[ 4272.011710] print_req_error: I/O error, dev sdc, sector 798740992&lt;BR /&gt;[ 4272.011713] BTRFS error (device sda): bdev /dev/sdc errs: wr 2, rd 0, flush 0, corrupt 0, gen 0&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Some additional information that might be useful:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;I started out with the Yocto state corresponding to LSDK-19.03 (as per yocto_2.7 tag in the qoriq-components/yocto-sdk repo). However this included the linux-qoriq kernel 4.19.26+gc0c214110624 which crashed at boot in ahci-qoriq driver's fsl_sata_errata_379364() function. That was fixed by upgrading to kernel 4.19.46+g1a4cab2c597d from LSDK-19.06. (some other components such as u-boot and management-complex firmware had to be upgraded to LSDK-19.06 state too to make the new kernel work)&lt;/LI&gt;&lt;LI&gt;I did not see such problems during disk I/O with only 4 SSDs connected directly to the board's SATA connectors. Also no problems with 4 SSDs connected to only 1 PCIe card, in either slot, as opposed to using 2 PCIe cards.&lt;/LI&gt;&lt;LI&gt;The same problems appeared when using 2 PCIe cards with 2 SSDs each (total 4 SSDs instead of 8).&lt;/LI&gt;&lt;LI&gt;We tried powering half of the SSDs with an external power supply, in case the internal one was too weak for 8 SSDs, but it doesn't seem to make a difference - errors appeared either way.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Does anyone have an idea what the problem could be? I'm pretty much lost at this point. It would be really nice to get the 8 SSDs to work reliably, but so far, no luck. Probably the way forward is to do more tests to rule out hardware problems with the PCIe cards or SSDs themselves.&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;But I don't understand why doing disk I/O would cause mmc1 timeouts. Yes, I'm booting from SD card, but that's mmc0 - mmc1 is (as far as I know) the internal eMMC which isn't being used at all (when mounting the btrfs raid0, only /dev/sd[a-h] are used).&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;Thanks in advance for any hints.&lt;/P&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Wed, 17 Jul 2019 12:08:10 GMT</pubDate>
      <guid>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/931124#M4474</guid>
      <dc:creator>daniel_klauer</dc:creator>
      <dc:date>2019-07-17T12:08:10Z</dc:date>
    </item>
    <item>
      <title>Re: lx2160ardb: kernel reports problems during heavy disk I/O with PCIe/SATA adapters</title>
      <link>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/931125#M4475</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;P&gt;Hello &lt;A _jive_internal="true" data-content-finding="Community" data-userid="344168" data-username="daniel.klauer@gin.de" href="https://community.nxp.com/people/daniel.klauer@gin.de"&gt;&lt;SPAN style="color: #0066cc; text-decoration: underline; "&gt;Daniel Klauer&lt;/SPAN&gt;&lt;/A&gt;,&lt;/P&gt;&lt;P&gt;&lt;/P&gt;&lt;P style="margin: 0cm 0cm 0pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;This issue could be caused by many reasons. It should be related to NCQ.&lt;/SPAN&gt;&lt;/P&gt;&lt;P style="margin: 0cm 0cm 0pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;It is more likely to happen when system is under heavy load.&lt;/SPAN&gt;&lt;/P&gt;&lt;P style="margin: 0cm 0cm 0pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;You can try the followings:&lt;/SPAN&gt;&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Try another disk to check if it was caused by disk. Different disk may behave differently. &lt;/SPAN&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Some disk even doesn&lt;/SPAN&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;’&lt;SPAN&gt;t support NCQ.&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI style="color: #000000; font-size: 10.5pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Disable NCQ by: &lt;STRONG&gt;echo 1 &amp;gt; /sys/block/&lt;SPAN style="color: red;"&gt;sda&lt;/SPAN&gt;/device/queue_depth&lt;/STRONG&gt; or &lt;/SPAN&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Here sda may need change according to your system. &lt;/SPAN&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Pass &lt;/SPAN&gt;&lt;STRONG style="font-size: 10.5pt; "&gt;“&lt;SPAN&gt;libata.force=noncq&lt;/SPAN&gt;”&lt;/STRONG&gt;&lt;SPAN style="font-size: 10.5pt;"&gt; to your bootargs. &lt;/SPAN&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P style="margin: 0cm 0cm 0pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P style="margin: 0cm 0cm 0pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;BTW: &lt;/SPAN&gt;&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;If NCQ was disabled, the performance will be impacted.&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI style="color: #000000; font-size: 10.5pt;"&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Even though such error occurred, there is no data lost on disk according to my test.&lt;/SPAN&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Thanks,&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN style="font-size: 10.5pt;"&gt;Yiping&lt;/SPAN&gt;&lt;/P&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Fri, 27 Sep 2019 07:50:29 GMT</pubDate>
      <guid>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/931125#M4475</guid>
      <dc:creator>yipingwang</dc:creator>
      <dc:date>2019-09-27T07:50:29Z</dc:date>
    </item>
    <item>
      <title>Re: lx2160ardb: kernel reports problems during heavy disk I/O with PCIe/SATA adapters</title>
      <link>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/1533989#M11254</link>
      <description>&lt;P&gt;Hi all, was this suggestion useful to solve the issue?&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;I've a similar issue with a LS1046 and mSATA SSD, where we face sporadic I/O Error on SATA bus traced as follows:&lt;/P&gt;&lt;P&gt;---&lt;/P&gt;&lt;P&gt;Sep 27 12:22:01 OTN kernel: [3109812.437650] sd 0:0:0:0: [sda] tag#28 UNKNOWN(0x2003) Result: hostbyte=0x00 driverbyte=0x06&lt;BR /&gt;Sep 27 12:22:01 OTN kernel: [3109812.437664] sd 0:0:0:0: [sda] tag#28 CDB: opcode=0x2a 2a 00 01 50 a1 cd 00 0a 00 00&lt;BR /&gt;Sep 27 12:22:01 OTN kernel: [3109812.437676] blk_update_request: I/O error, dev sda, sector 22061517 op 0x1:(WRITE) flags 0x4000 phys_seg 46 prio class 0&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;....&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;that does not match with SSD status. There are no badblocks or error on disk.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Do you have any suggestion?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Thanks,&lt;/P&gt;&lt;P&gt;Michele&lt;/P&gt;</description>
      <pubDate>Fri, 07 Oct 2022 15:49:23 GMT</pubDate>
      <guid>https://community.nxp.com/t5/Layerscape/lx2160ardb-kernel-reports-problems-during-heavy-disk-I-O-with/m-p/1533989#M11254</guid>
      <dc:creator>mdecandia</dc:creator>
      <dc:date>2022-10-07T15:49:23Z</dc:date>
    </item>
  </channel>
</rss>

