Several gpu crashes/invalid output, when running opengl es cts Hi, For over three years we have been shipping a device on i.MX6QuadPlus, built on Yocto hardknott with Qt 6.3.2 on Weston. As the product grew we started getting more and more crash reports from users, and we could not find a cause in our own code or in Qt. Some of it we did fix - by retuning the DDR timings, disabling shadows in certain views, increasing some internal buffer sizes, and by running the app with GPU_VIV_EXT_RESOLVE=0 - but we still get reports, and a good part of them are GPU related and show up in the log like this: kernel: *** GPU DRV CONFIG *** kernel: Galcore version 6.4.3.336687 kernel: Galcore options: ... kernel: [galcore]: Stop driver to keep scene. That is the driver's own hang report - the monitor timer sees no progress, gckKERNEL_Recovery dumps the GPU state, and since recovery=0 the driver stops serving instead of resetting the core(recovery=1 is not an options for us, due needing to restart every gui application). The screen is frozen for good, Weston often cannot be killed even with SIGKILL, and the only way out is a power cut - which is quite bad for our customers. Since we had already tried a great many things at the Qt and application level and only got workarounds out of it, I decided to stop testing high-level functionality and test the OpenGL ES entry points directly instead - if they behave correctly the fault is ours, and if they do not then it is the driver/hardware fault(at least partial). I did that with VK-GL-CTS (https://github.com/KhronosGroup/VK-GL-CTS), first ported to Rust to make cross-compiling easy(still in progress, so only about half of the relevant cases could be checked so far). Case names are identical to upstream deqp-gles2/gles3/gles31, and everything passes on Mesa llvmpipe and on casual laptop on Ubuntu 24.04 with also mesa driver, so a failure on the board is a statement about the board. I ran it on the hardware over a weekend and with help of AI I was able to find and minimise 14 distinct defects, each now a standalone reproducer: a plain Rust project with no dependencies, cargo build and nothing else, GLSL in its own file. (just run - runs locally, you can extract just arm parts, into only building step, to gather binary(you need to install first cargo install --locked cargo-zigbuild) Problems still present on 6.4.11.p4, the newest driver we can build for this part: 0001 GPU lockup - synchronization.inter_invocation.ssbo_atomic_read_write, alone, on a freshly booted board. Only a reboot clears it, and the process is unkillable so the reboot itself takes 8-10 minutes. 0002 GPU lockup - synchronization.inter_invocation.ssbo_atomic_overwrite, on its own. 0003 GPU lockup - ten of the twenty synchronization.inter_invocation.* cases hang alone, the other ten do not, so it is a boundary and not "compute is broken". All the atomic ssbo/image variants. All twenty swept ten runs each, every run from its own reboot. 0010 Link failure - 7 ubo cases where both stages read 47 members of the same std140 block. Either stage alone links, together they do not, glGetProgramInfoLog is empty, and nothing is near a limit the driver itself reports (3 blocks per stage vs 16, 704 bytes vs 65536). 47 reads is the least that does it, 46 links. 0011 Wrong result - shaders.invariance.highp.loop_*: two shaders computing invariant gl_Position from the same expression disagree by a few pixels of depth. 0013 Wrong status enum - fbo.completeness.size.distinct: a context requested as ES 2.0 reports ES 3.1, then answers completeness by the ES 2.0 rule and returns an enum ES 3.x does not define. Fixed by the 6.4.3.p2 -> 6.4.11.p4 jump: 0004 GPU lockup - image_load_store.cube.qualifiers.*_r32f (the r32ui/r32i ones beside them were fine) 0006 GPU lockup - image_load_store.* whenever the image is layered, 8 of 21 layered vs 0 of 7 over 2d, intermittent 0005 Client freeze - compute.indirect_dispatch.gen_in_compute.empty_command: glMapBufferRange never returns 0014 Client freeze - the same, via upload_buffer.empty_command, once a dispatch has already been mapped 0007 Wrong result - a chain of eighteen && in a compute shader that also declares an atomic counter comes out false when every term is true (44 of 2007 ssbo.layout cases) 0008 Compiler - the reserved-word table is for the wrong language version, both ways 0009 Wrong result - vec3 == vec3 false for two equal vectors, for a struct member returned through an inout parameter 0012 Compiler - mediump vec2(1.0, 1.0) compiles, where the ESSL 1.00 grammar has no place for a precision qualifier Tested on same board both times: i.MX6QP silicon rev 1.0, 2 GiB DDR, LVDS 1280x1024@60, Weston on fbdev with use-g2d=1, GL_RENDERER "Vivante GC2000+". old: hardknott, BSP imx-5.10.52-2.1.0, kernel 5.10.52, galcore 6.4.3.p2.336687, imx-gpu-viv 1:6.4.3.p2.2-aarch32, Weston 9.0.0.imx, Qt 6.3.2 new: wrynose, BSP imx-6.18.20-2.0.0, kernel 6.18.20, galcore 6.4.11.p4.1190909, imx-gpu-viv 1:6.4.11.p4.6-aarch32, Weston 10.0.5.imx, Qt 6.11.0 CONFIG_MXC_GPU_VIV=y, recovery=0 and stuckDump=0 on both; the timeout went 20000 -> 30000 ms and 6.4.11.p4 adds softReset=1. So the version bump helps, but some problems survived. We will be moving to i.MX8 soon because of i.MX6 availability, but our installed base keeps the i.MX6 hardware either way, so it would be good to see such problems fixed in future Is there any chance these get fixed in a newer version of the driver, even if that forces us to bump the Yocto version? I am afraid this class of problem may to some extent be present on i.MX8 as well, so has any testing with the Vulkan/OpenGL CTS been done, is it being done now, or is it planned? Because without passing that(plus some fuzzing like random executing of cts functions), I do not think our random crashes are fixable in the layers above the driver Re: Several gpu crashes/invalid output, when running opengl es cts Hello,
The latest publicly released driver for i.MX 6/7 is imx-gpu-viv 6.4.11.p4.6 , shipped with the wrynose BSP (imx-6.18.20-2.0.0). The release notes describe that jump as bringing "bug fixes, performance optimizations" for the i.MX 6/7/8 line, which matches your observation that 8 of 14 defects were resolved between .p2 and .p4 . And yes, it is being done — but with important caveats by platform. For i.MX8 with Vivante (VSI) GPU — CTS is running. The internal Linux Factory test pipeline runs both opengl-es-cts and vulkan-cts packages against i.MX8 boards as part of every release candidate cycle. Defects found are tracked in the Linux Factory Jira project CTS on i.MX8M Nano, i.MX95. This means the Vivante GC7000-series GPUs in i.MX8M Plus, i.MX8QuadMax etc. do go through systematic conformance testing before each GA release.
For i.MX9 with Mali / OSS Mesa
The release notes explicitly state for i.MX 95/952 with the Mesa OSS GPU stack: "OpenGL ES11, Vulkan 1.4.5, and OpenCL 3.0 basic features are working, but conformance tests are not passed."The Mali DDK default path on i.MX9 does pass CTS, but the open-source Panfrost/PanVK path is still being brought to conformance.
For i.MX6 (GC2000+) specifically — CTS coverage is minimal
No internal evidence was found of systematic deqp/CTS runs against the GC2000+ in the current Linux Factory pipeline. The GC2000+ only supports OpenGL ES 3.0 (not 3.1/3.2), and the test infrastructure appears to target the newer i.MX8/9 boards. Your work with VK-GL-CTS is the most thorough conformance-level testing of this specific IP that is visible in any internal channel.
The bottom line: a future i.MX6 driver patch is not guaranteed, but providing the standalone reproducers via your open support thread is the correct path. For your i.MX8 migration, the conformance situation for the Vivante GC7000-series is materially better than GC2000+, and systematic CTS testing is part of the release process — though even there, active driver bugs in galcore 6.4.11.p4 are still being discovered and filed.
Regards Re: Several gpu crashes/invalid output, when running opengl es cts I don't see an attachment, so looks that I forgot to add it, so I add it again
View full article