2417321_en-US

取消
显示结果 
显示  仅  | 搜索替代 
您的意思是: 

2417321_en-US

2417321_en-US

The software encountered a "Synchronous Abort" error during startup: "handler, esr 0x02000000".

1. An intermittent fault (occurring once out of thousands of reboots): The software freezes during the u-boot process. The fault message is: "Synchronous Abort" handler, esr 0x02000000. Disassembling the code reveals the following call path: initr_pci → pci_init → dm_pciauto_config_device (pci_auto.c:371) →dm_pci_hose_probe_bus (pci-uclass.c:607) → dm_pciauto_prescan_setup_bridge→ When fetching pointer 0x8202fc20 after `bl dm_pci_get_bdf` and `ret`, an exception occurred. The problem is that `ret` was accessing an incorrect address, while the `asr` call is just a register shift and has no inherent error-prone parts (it does not access memory or jump).

2. The software uses CPLD to actively feed the watchdog (external watchdog). After the problem occurred, the external watchdog was reset using PORESET, but it failed.

Re: 软件启动时出现Synchronous Abort" handler, esr 0x02000000

Hello,

This is not likely an error in the asr instruction itself. The key symptom is that ret loads the PC from a corrupted or incorrect link register, causing instruction fetch from 0x8202fc20. The abort is therefore probably detected at instruction fetch, while the corruption occurred earlier.

Most likely causes

Investigate these areas first:

  1. Stack corruption or stack overflow

    • Check sp at every PCI recursion level.
    • Add stack canaries around the U-Boot stack.
    • Verify that PCI enumeration does not exceed the configured stack size.
  2. Invalid function pointer or corrupted LR

    • Capture x30/LR, SP, PC, and all registers immediately before and after bl dm_pci_get_bdf.
    • Confirm that the saved LR on the stack matches the expected return address.
    • Disassemble the exact binary, not only the source code.
  3. Memory corruption

    • Check DDR stability, ECC/error status, cache configuration, and DMA activity.
    • Temporarily disable data cache, speculative accesses, and concurrent DMA.
    • Fill unused memory and stack with known patterns, then inspect corruption after failure.
  4. PCIe configuration access timeout/error

    • Confirm that link training and controller reset are complete before enumeration.
    • Verify that absent devices return a legal completion error rather than hanging or producing invalid data.
    • Test with PCI initialization disabled; if the issue disappears, focus on PCIe hardware/configuration timing.

Why the external watchdog reset may fail

The watchdog can still be fed by a CPLD while the CPU is executing corrupted code, so the system may never reach the watchdog timeout. Also, PORESET may not reset the logic that is holding the system in the failed state, or its pulse may not meet the required width/sequence.

NXP documentation describes external watchdog logic as an independent reset trigger, but its actual reset path and affected domains must be verified in the SoC and board reset architecture.

Check with an oscilloscope or logic analyzer:

  • CPLD watchdog timeout output
  • PORESET at the SoC pin
  • HRESET/system reset
  • PMIC reset input/output
  • power rails during the reset attempt
  • reset-cause registers after recovery

Recommended debug experiment

Modify the watchdog policy so that:

  • U-Boot feeds the external watchdog only from a controlled periodic path.
  • The watchdog is not fed during PCI enumeration.
  • On timeout, the CPLD generates a reset pulse long enough for the SoC/PMIC specification.
  • Record reset cause and preserve the failing registers in SRAM or another retention area.

The most valuable next artifact is a complete exception dump containing PC, LR/x30, SP, ESR, FAR, SPSR, and the stack contents around the saved return address. Without that data, the failure location identifies where the corrupted return is detected, not where the corruption originated.

 

Regards

标记 (1)
无评分
版本历史
最后更新:
星期四
更新人: