Skip to content

feat: SG-42200: 10-bit Metal presentation - #1327

Draft
cedrik-fuoco-adsk wants to merge 65 commits into
AcademySoftwareFoundation:mainfrom
cedrik-fuoco-adsk:metal-10bit-macos
Draft

cedrik-fuoco-adsk wants to merge 65 commits into
AcademySoftwareFoundation:mainfrom
cedrik-fuoco-adsk:metal-10bit-macos

Conversation

@cedrik-fuoco-adsk

Copy link
Copy Markdown
Contributor

feat: SG-42200: 10-bit Metal presentation

Linked issues

Summarize your change

Adds a macOS-only 10-bit Metal presentation path for OpenRV when the user requests 10+2 display format and the hardware supports it. IPCore still renders through an offscreen OpenGL 2.1 context (GL-on-Metal); the new backend delivers pixels to the screen via IOSurface → CALayer instead of a QOpenGLWidget swap chain.

The work is built on a platform-neutral view abstraction that decouples RvDocument from GLView: shared code routes through viewWidget() and the session's control video device, so alternative backends (this Metal path, and the existing Vulkan 10-bit path on Linux) can plug in without touching every call site.

When 10-bit Metal is unavailable or fails at runtime, the app falls back to the legacy OpenGL GLView with logging.

Describe the reason for the change

OpenRV's main window has historically been a QOpenGLWidget. True 10-bit display on macOS is unreliable through that path: requesting a 10+2 GL surface format does not reliably produce correct scanout, and a QOpenGLWidget anywhere in the window hierarchy can force Qt onto _NSOpenGLViewBackingLayer, breaking CALayer/IOSurface presentation (including 4× tiling artifacts with device pixel ratio scaling).

This PR introduces a dedicated MetalView (native NSView + CALayer backed by IOSurface) and QTMetalVideoDevice, which wraps it as a TwkGLF::GLVideoDevice so ImageRenderer and the rest of IPCore continue to render into offscreen FBOs unchanged. The zero-copy path blits RGBA16F → IOSurface-backed RGB10_A2; when GL–IOSurface interop is unavailable, a GPU-blit + packed readback fallback avoids the previous per-pixel CPU pack loop (~25M float ops/frame at 4K).

Additional fixes address real usability bugs on the new path: black flash on window re-expose (alt-tab / uncover), context/thread safety for upload threads, interop failure retry instead of permanently disabling zero-copy, and diagnostics overlay integration via Qt::AA_ShareOpenGLContexts.

Describe what you have tested and on which operating system

  • macOS — Apple Silicon, 10-bit-capable display, USE_METAL=ON, 10+2 display prefs: main window presents correctly, no banding vs 8-bit GL path
  • macOS — Alt-tab / window uncover: no black flash; last frame re-presented immediately
  • macOS — Resize, Live Review panel, activity timer / playback scrubbing
  • macOS — Fallback when supports10BitPresentation() fails or GL context creation fails → GLView with log message
  • macOS — CPU fallback path when IOSurface GL interop unavailable (performance acceptable vs old float readback loop)

OS tested: macOS (version / hardware — please confirm)

Add a list of changes, and note any that might need special attention during the review

Core architecture (review carefully)

Area Change
View abstraction RvDocument uses QWidget* m_viewWidget + viewWidget() instead of hardcoded m_glView; MuUICommands / PyUICommands updated for null-safe coordinate mapping
RvApplication Tolerates null view(); primary display group driven by session control video device
MetalView QWidget with WA_PaintOnScreen / WA_OpaquePaintEvent, paintEngine() → nullptr, IOSurface cache + re-present on expose, coalesced UpdateRequest, renderImmediately() for synchronous redraw
QTMetalVideoDevice Offscreen GL render FBO → IOSurface ring buffer (zero-copy) or GPU-blit flip FBO + packed glReadPixels (CPU fallback); interop failure latch with periodic retry
RvDocument Selects MetalView vs GLView when 10+2 prefs + MetalView::supports10BitPresentation(); fallbackMetalToGLView() on context failure
ImageRenderer / Session Upload-thread device sharing, synchronous upload fallback, explicit GL flush

Known limitations (not blockers, but reviewers should be aware)

  • Metal activates only when 10+2 display prefs are set and hardware check passes; changing prefs requires a new window
  • Hardware stereo (shutter glasses) not supported on Metal; software stereo modes should work
  • VSync, double-buffer, GL pixel format prefs are no-ops while on Metal (apply after GL fallback)
  • Presentation Mode secondary displays remain 8-bit (separate DesktopVideoDevice path, not ported)

cedrik-fuoco-adsk and others added 13 commits June 10, 2026 14:58
Decouple RvDocument from GLView by routing the active view through a
generic QWidget* m_viewWidget and a viewWidget() accessor, instead of
hardcoding m_glView at every shared call site (focus, geometry,
stacked-layout, popup/menu coordinate mapping in Mu/Py UI commands).

RvApplication now tolerates a null view()/GL context and uses the
session's control video device for the primary display group, so an
alternative presentation backend can be plugged in.

No backend-specific code is introduced here; the GL path behaves exactly
as before. This is the shared base for the Metal (macOS) and Vulkan
(Linux) 10-bit presentation backends.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Builds on the neutral view-abstraction base. Adds VulkanView (a Vulkan
swapchain over a 30-bit X11/Wayland visual) and QTVulkanVideoDevice
(offscreen 16F FBO + GL<->Vulkan zero-copy interop), wired into
RvDocument behind '#if defined(PLATFORM_LINUX)'.

Backend selection is driven by the display-depth preference: a 10-bit
request (RGB10 + A2) routes to Vulkan when VulkanView::supports10Bit
Presentation() is true, otherwise it falls back to the OpenGL GLView.
On the Vulkan path m_glView is null, so DiagnosticsView (a QOpenGLWidget)
is not created.

Vulkan (headers + loader) is fetched as a managed Linux dependency
(cmake/dependencies/vulkan.cmake) and linked unconditionally on Linux.
The shared view plumbing lives in the neutral base; this commit only
adds the Vulkan-specific pieces.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Replaced platform-specific size policy and content size handling with unified methods for setting active view sizes.
- Introduced new methods: setActiveViewContentSize, setActiveViewMinimumContentSize, and activeViewFirstPaintCompleted to enhance code clarity and maintainability.
- Updated resize logic to utilize the new methods, ensuring consistent behavior across platforms.

This refactor simplifies the management of view sizes and improves the overall structure of the RvDocument class.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Added QTVulkanVideoDevice, a new class that serves as a GL-to-Vulkan bridge, enabling 10-bit display on Linux by transferring frames from an OpenGL rendering pipeline to a Vulkan swapchain.
- Implemented lifecycle management for VulkanView, including lazy GL context setup and buffer synchronization.
- Enhanced video device API to support Vulkan, ensuring compatibility with existing GLVideoDevice functionality.
- Updated main application to ensure all QOpenGLContexts share resources, preventing issues with font texture uploads.

This addition significantly improves rendering capabilities on Linux platforms, leveraging Vulkan's advanced features.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…ew and QTVulkanVideoDevice

- Added error handling for Vulkan instance creation and memory allocation failures, ensuring proper cleanup and resource management.
- Improved format validation logic to ensure only supported formats are used, with detailed warnings for unsupported configurations.
- Updated comments for clarity regarding format assumptions and fallback mechanisms.

These changes enhance the robustness of the Vulkan integration, improving error reporting and preventing potential crashes due to unsupported formats.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Updated CMake configuration to fetch Vulkan dependencies for both Linux and Windows, ensuring consistent availability across platforms.
- Improved error handling and format validation in VulkanView and QTVulkanVideoDevice, enhancing robustness and preventing crashes due to unsupported formats.
- Added support for 10-bit presentation on Windows, aligning with existing Linux functionality.
- Refactored Vulkan-related code to streamline platform-specific implementations and improve maintainability.

These changes significantly enhance Vulkan integration, providing better error reporting and expanding functionality for Windows users

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Added a fallback mechanism in RvDocument to switch from VulkanView to GLView upon runtime failures, enhancing stability on Linux and Windows platforms.
- Introduced error handling in VulkanView to request a fallback when Vulkan initialization fails or when the swapchain becomes out of date.
- Updated RvDocument destructor to ensure proper cleanup of Vulkan resources during closure.
- Enhanced VulkanView with methods to manage presentation conditions and handle swapchain recreation.

These changes improve the robustness of the Vulkan integration, ensuring a seamless user experience by gracefully handling failures and maintaining functionality across platforms.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…r insights during development and troubleshooting.

These changes streamline the debugging process and improve code readability, ensuring a more maintainable codebase.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Added QTVulkanVideoDevice class to facilitate rendering from OpenGL to Vulkan, enabling true 10-bit display on Linux.
- Implemented lifecycle management, including lazy GL context setup and buffer synchronization between GL and Vulkan.
- Enhanced VulkanView to support both A2B10G10R10 and A2R10G10B10 formats for improved compatibility.
- Updated logging for better insights into GPU operations and format handling.

These changes significantly enhance Vulkan integration, providing a robust solution for high-quality video rendering across platforms.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
@cedrik-fuoco-adsk
cedrik-fuoco-adsk force-pushed the metal-10bit-macos branch 2 times, most recently from f4e483d to 4c71ed7 Compare July 6, 2026 14:55
cedrik-fuoco-adsk and others added 9 commits July 6, 2026 11:23
…lkan swapchain (VulkanView) fed by a GL↔Vulkan shared image (QTVulkanVideoDevice). That present was fully synchronous and single-frame-in-flight. Every frame ended in a blocking vkWaitForFences under FIFO vsync, and every resize step tore down and re-exported the shared image and recreated the swapchain. This saturated the Qt event loop, so dragging a dock splitter over the media view was choppy and playback would stall during the drag then fast-forward on release.

This change removes that cost in three steps:

- Per-frame sync objects and the GL interop resources become per-in-flight-slot rings.
- The present pipelines to 2 frames in flight, with the fence wait moved from the end of the frame to the start (the throttle is now FIFO back-pressure plus that wait). The present-wait semaphore is tied to the swapchain image, an imagesInFlight fence map avoids semaphore reuse, and a small drain submit keeps the GL/VK semaphore pair balanced when a frame is dropped on OUT_OF_DATE.
- The shared image is allocated grow-only to a capacity, and only the used region is transferred, so a resize within bounds no longer re-exports it.

Format and channel handling, the OUT_OF_DATE-only swapchain recreate policy, and the CPU fallback are all preserved. The result: resize tracks the cursor and playback cadence stays locked to refresh, matching the OpenGL GLView path.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
- Introduced a CPU fallback mechanism in QTVulkanVideoDevice to handle scenarios where GPU interop is unavailable.
- Added methods to ensure and clean up CPU fallback targets, allowing for efficient readback of pixel data.
- Enhanced the present function to utilize the CPU fallback, ensuring compatibility across platforms when Vulkan interop fails.

These changes improve the robustness of the Vulkan integration, providing a reliable alternative for rendering when GPU resources are not accessible.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…on path

- Probe the driver once per device for an exportable shared-image format,
  tiling, and dedicated-allocation requirements instead of relying on
  vendor/platform assumptions.
- Add RV_VULKAN_FORCE_CPU_PRESENT, RV_VULKAN_FORCE_TILING, and
  RV_VULKAN_FORCE_NO_DEDICATED environment overrides, applied consistently
  to both the Vulkan export and the GL import.
- Record the resolved presentation path (zero-copy, CPU readback, or OpenGL)
  and the reason for the choice in a single per-session startup log.
- Latch GL-side interop failures so the session stays on the CPU fallback
  instead of retrying the failing import every frame.
- Require glMemoryObjectParameterivEXT on Windows so dedicated Vulkan
  exports are imported as dedicated memory objects.
- Request a Vulkan 1.1 instance and fall back to the loader default if it
  is unavailable.
- Pass the actual tiling and dedicated-allocation flags through
  SharedImageInfo so the GL import matches the Vulkan export exactly.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Resolve the conflicts with the viewport refactor that landed in AcademySoftwareFoundation#1375,
which split GLView into a plain-QWidget facade plus a native GLWindow
(QOpenGLWindow).

src/bin/apps/rv/main.cpp
  Both sides add Qt::AA_ShareOpenGLContexts. Keep main's block, since it
  also carries the Windows QQuickWindow::setGraphicsApi() fix, and fold
  this branch's Vulkan rationale into its comment.

src/lib/app/RvCommon/GLView.cpp
  Take main's version verbatim. The only changes this branch had were the
  Linux "-debug gpu" diagnostics, which now belong in GLWindow.cpp: it
  owns initializeGL() and is the sole caller of GLView::rvGLFormat(), so
  the formatSummary()/envOrUnset() helpers live in one translation unit
  rather than being duplicated across two.

src/lib/app/RvCommon/RvDocument.cpp
  initializeSession() picked up "if (!m_glView) return;" plus an
  unconditional m_glView->makeCurrent() from main. On the Vulkan
  presentation path m_glView is null by design, so that early return
  would have skipped session creation altogether. Guard on m_viewWidget
  instead, and only makeCurrent() when a GLView is present -- VulkanView
  already makes its own GL context current before calling in.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
The -debug gpu surface-format and GL vendor/renderer/version baseline was
added while this branch was Linux-only. Windows is now a supported 10-bit
presentation target, and is where the GL/Vulkan interop surprises have
been, so emit the same baseline there.

The XDG_SESSION_TYPE / WAYLAND_DISPLAY / DISPLAY line stays Linux-only, as
does envOrUnset(), which would otherwise be an unused function elsewhere.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
@cedrik-fuoco-adsk
cedrik-fuoco-adsk force-pushed the metal-10bit-macos branch 6 times, most recently from e8b78e3 to d584486 Compare September 28, 2026 19:16
…rmat

QTVulkanVideoDevice allocated its offscreen FBO as GL_RGBA16F_ARB for
every device. The control viewport needs that depth -- session->render()
composites the whole main view into it across blended passes -- but a
passive presentation output never does. It is only ever a blit
destination for the inherited transfer()/transfer2()/fillWithTexture(),
which hand over an already-composited frame.

Keeping it at 16F there cost two full passes' worth of bandwidth at a
3840x2160 output: transfer() wrote 8 bytes/px, some 66MB, and
syncBuffers() read all of it back to convert down to the 10-bit shared
image. Matching the shared image's format halves both, and the
conversion happens once, in a blit that was already going to run. The
output is 10-bit either way, so no precision is lost that the present
did not already discard.

Measured on a 4K presentation output: the per-frame GPU wait dropped
from 23ms to 5.5-16ms, frame interval from 33-45ms to 18-29ms, and
eventToRetire from 38-49ms to 22-33ms -- parity with presentation off.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
The OpenGL presentation path never lets the second display gate the
viewport's loop: DesktopVideoDevice::syncBuffers() is a coalesced
QOpenGLWidget::update() on an empty paintGL() that Qt drops when it
falls behind. The Vulkan path queued its output work unconditionally,
every frame, and at a 4K output that work is what the control viewport
ends up waiting for in vkWaitForFences.

Gate the output present on whether this device's own GPU work has caught
up, rather than on whether the swapchain is full -- the latter never
fires, because a loop slower than the display always leaves the queue
room. canPresentNow() is checked from syncBuffers() ahead of any GL
work, so a skipped frame costs nothing at all instead of costing the
full-resolution blit, and no GL semaphore has been signalled yet so
there is nothing to rebalance. The blocking paths inside
presentSharedImage()/presentPixelData() keep a second, later skip for
safety, which does have to drain the semaphore pair.

A skipped present must be retried or the output is left on the frame
before the one just composited, with nothing else coming back for it:
the control viewport only renders when the session asks. Qt's update()
is a dirty flag it must eventually honour, so parity needs the same
guarantee -- requestBestEffortRetry() posts a coalesced UpdateRequest
that render() serves from its passive-output branch, and a staleness
timer forces a blocking frame through if the output has gone unpresented
for more than 100ms.

Note this does not currently fire on the hardware it was developed
against: with a one-deep control pipeline the viewport waits for its own
GPU work each frame, which gives the output's fences time to retire, so
the gate passes every time and outputPresent never reaches zero. It is
kept as the protection for the reverse balance -- a slower output
display, a heavier output scene, or a faster control GPU -- which also
means its skip and starvation paths are untested in practice. Isolated
in its own commit so it can be dropped with a single revert.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…nd swap

Switching between 8- and 10-bit reset the display transfer function from
sRGB to None, discarding any assigned display profile with it.

Two causes, in sequence.

IPGraph::deviceChanged() adopted newDevice->physicalDevice()
unconditionally. VideoDevice's constructor seeds m_physicalDevice with
the device itself, and a viewport device only learns the monitor it sits
on when it first renders, via setAbsolutePosition() ->
deviceFromPosition(). Its one caller is
Session::setControlVideoDevice(), which during a backend swap runs on a
view that has never rendered -- so physicalDevice() was still that view,
and the display group's device.name was rewritten from the monitor
("Dell Inc. DELL U2725QE DP-1") to the viewport's own name ("RV Main
Window (Vulkan)/0x..."). Only adopt a physical device the new device
actually knows; the monitor has not changed, only the object drawing to
it.

RvApplication::rebuildDesktopVideoDevices() then called
IPGraph::setPhysicalDevices() to refresh stale device pointers, but that
is the startup routine: it deletes every DisplayGroupIPNode and rebuilds
them, and a new display group comes with a new colorPipeline holding
default contents. So the group was discarded and its colour state with
it. Add IPGraph::refreshPhysicalDevices(), which re-points existing
groups at the rebuilt devices by (module name, device name) -- the same
key display profiles are stored under, and stable across a rebuild
because createDesktopVideoDevices() names every screen from its QScreen
regardless of backend. Groups are created or deleted only for devices
that genuinely appeared or vanished.

It also clears a group's output device when that pointer is not among
the new devices. findDisplayGroupByDevice() compares raw pointers, so a
dangling one can alias a freshly allocated device at the same address
and return the wrong group. The control device is kept, being alive and
never in the module list.

Both changes are needed: the guard alone still lost the group to the
rebuild, and the refresh alone could not match a group whose name had
already been rewritten.

Note this touches IPCore paths shared with SDI/AJA output and
multi-monitor setups, which were not exercised here. The guard assumes a
device reporting itself as its own physical device carries no monitor
information -- true for viewport devices, and deviceChanged() only ever
runs on the control device today, but a real physical device
legitimately is its own physical device.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
QTVulkanVideoDevice now compares the GL physical-device UUID against the
Vulkan physical device via VulkanWindow::physicalDeviceMatchesUUID().
If the UUIDs do not match (or cannot be queried), it falls back to the
CPU pack-and-upload path instead of exporting GL memory to a different
GPU. The result is cached and can be reset.

VulkanWindow now creates a Vulkan 1.1 instance so
vkGetPhysicalDeviceProperties2 and device UUIDs are available, and
tightens physical-device selection to require a graphics+present queue
family, VK_KHR_swapchain, and a 10-bit surface format when applicable.

RvDocument's Vulkan-to-OpenGL fallback now preserves the requested display
depth except when recovering from a 10-bit Vulkan failure, in which case
it explicitly falls back to 8-bit OpenGL with clearer log messages.

MuUICommands::colorAtCursor now reads the cursor pixel through the active
GLVideoDevice with glReadPixels instead of requiring a GLView and
QImage, making it backend-agnostic.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
With no context current, glGetError() on Windows returns
GL_INVALID_OPERATION for every call, for as long as nothing is current.
TWK_GLDEBUG reported that as a GL error at each instrumented site that
followed, so one missing context surfaced as a dozen copies of itself,
attributed to whichever innocent line checked next. A missing context
during presentation teardown showed up as an error inside makeCurrent(),
a frame late and in the wrong place.

twkGlAnyContextIsCurrent() answers the question directly.
QOpenGLContext::currentContext() only knows about contexts Qt made
current, and TwkGLFFBO's FBOVideoDevice binds its own natively, so
trusting Qt alone would claim "no context" while a perfectly good one is
current and would suppress the real errors the macro exists to print.
glGetString() settles the cases Qt cannot see, and is only reached when
Qt says no.

twkGlPrintError() now checks that first and, when nothing is current,
reports once per episode at the first site to notice, resetting when a
context returns so a later episode is not swallowed.

Both are declared outside the NDEBUG guard. TWK_GLDEBUG still compiles
out in release -- polling glGetError() at every instrumented site is a
debug-only cost -- but "no current GL context" is not instrumentation.
It fires only when GL work cannot land, which is a fault in a release
build too, and the callers that need to say so are compiled in both.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
glDeleteFramebuffers() and friends are silent no-ops with no context
current: the C++ object goes away, the driver's does not, and nothing
says so. Teardown paths are where this bites, because they run from
destructors and event callbacks rather than from inside a render, so
nothing has arranged a context for them.

Fixing that per call site does not converge -- each one found reveals
the next, because nothing in the code states the invariant. This scope
states it once. Open it at the top of anything that deletes GL objects
and the question stops being the caller's problem.

It resolves a context in three steps: already current, so do nothing --
the common case, a pointer compare, safe to put on paths that also run
mid-render; else makeCurrent() on a supplied device, which is cheaper
and is the context the objects were most likely created under; else a
process-lifetime fallback. Destruction restores what was current, which
-- because it only acquires when nothing was -- means making nothing
current again.

The fallback is what lets this work in destructors that have already
had their device pointers cleared out from under them, which is exactly
where the problem lives. It insists on QOpenGLContext::globalShareContext():
GL names belong to a share group, so deleting an FBO under a context
outside the group that created it does nothing at all, and a
non-sharing fallback would look like a fix while behaving like the bug.
It is created once and never destroyed, because the paths needing it
run while the application object is being torn down, so any owner
freeing it at static-destruction time would free it either too early to
be useful or after QGuiApplication has gone.

If no context can be resolved it reports once and does nothing. It
never throws and never aborts: a teardown helper must not be the reason
a process fails to exit.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Leaving presentation mode resizes the main view, which reaches
Session::deviceSizeChanged() from the resize rather than from a render,
so no context is current. It flushed the renderer's entire ImageFBO
pool into nothing: the C++ objects went away, the driver's did not, and
the only sign was a GL_ERROR from ~GLFBO after the fact.

Open a GLContextScope at the owner instead of at each caller.
ImageFBOManager::flushImageFBOs() and destroyImageFBO() cover every
path that reaches them, present and future, so the hand-placed
makeCurrent() in Session::clearVideoDeviceCaches() and the context
restore in ~RvDocument go away rather than accumulating. Session passes
the device it already knows -- deviceSizeChanged() its argument,
clearVideoDeviceCaches() the control device -- which is cheaper than
the fallback and is the context the objects were created under.

ImageRenderer needs the same at three more points.
Device::clearFBOs() deletes the FBO ring buffer, clearState() deletes
the program cache immediately after flushing the pool -- so the pool's
own scope closing would leave a gap -- and ~ImageRenderer covers both
plus the program cache object. Session clears the renderer's device
pointers before destroying it, so by then it has nothing to ask and the
scope falls back to its own context; that case is the reason the
fallback exists.

~GLFBO keeps a backstop. With no context current it reports once,
naming the destruction site, and skips the GL calls rather than
pretending the names were released. It asks only for an FBO that owns
GL names: the GLFBO(const GLVideoDevice*) constructor builds a handle
onto whatever the device has bound, with no id and no PBO, so
destroying one issues nothing and needs no context. Checking the
context before checking for work reported a leak that cannot happen on
the ordinary path where a device outlives its window.

It reports rather than asserts. A leaked FBO is worth a line of output;
it is not worth aborting a shutdown that would otherwise have
completed, least of all in the debug build someone is using to diagnose
that shutdown.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
UninitPBOPools() never released a PBO. PBOWrap::uninitPBOPool() only
flipped an initialized flag, leaving gPoolToGPU and gPoolFromGPU -- both
file-scope statics -- to delete their buffers from ~GLPixelBufferObjectPool
at static destruction, after main() has returned. Qt is gone by then and
no context can be obtained on any platform, so every glDeleteBuffers()
in there was a silent no-op and the driver reclaimed the memory with the
process.

Give the pool a clear() that does the release, and call it from
uninitPBOPool() so it happens while the application is still up.
clear() empties the containers and resets the accounting as well as
deleting, because the destructor still runs later and would otherwise
walk the same entries a second time.

UninitPBOPools() is called from main() once the event loop has returned,
so the views and their contexts are already gone and nothing is current.
A GLContextScope supplies the fallback context -- still available there,
since the QApplication outlives the call -- which also covers the
GLSyncObject fences deleted alongside each buffer.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
RvConsoleWindow installs its cout/cerr redirect under

    #if defined(NDEBUG) || defined(PLATFORM_WINDOWS)

and took it down under

    #if defined(NDEBUG) || !defined(PLATFORM_WINDOWS)

A Windows debug build is the one combination where those disagree, so
there the redirect went in and never came out. ConsoleBuf stayed on
cout and cerr with m_console pointing at the destroyed window.

main() deletes RvApplication before finalizePython(), and Py_Finalize's
garbage collection can still write -- a ResourceWarning from an
unclosed socket, for one. That write reached ConsoleBuf, followed
m_console into freed memory, and locked a QMutex whose bits happened to
read "contended". Nothing ever releases it, so RV hung on the way out
and had to be killed. It also meant any late shutdown output was lost.

Guard on having installed the redirect rather than on a second
attempt at the same #if. m_stdoutBuf and m_stderrBuf are non-null only
if the install ran, and processLastTextBuffer() nulls them if it got
there first, so this is correct in every build and safe to run twice.
Delete the ConsoleBuf too, so nothing is left pointing at the window.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
RV never calls quit(). It relies entirely on Qt's
quitOnLastWindowClosed, and QApplicationPrivate::shouldQuit() counts
every visible top-level widget carrying WA_QuitOnClose, which is on by
default. So any auxiliary window left visible when the session window
goes stops exec() from ever returning: the process stays up with a
stray dialog on screen and has to be killed.

The console reached that state on its own.
RvConsoleWindow::processTextBuffer() calls show() and raise() for any
line the show-on preference considers interesting, and at showOn=3
processLine() returns true for every line. It runs from a queued
event, so it lands after ~RvDocument has closed the console, reopening
it as the last visible window while the rest of shutdown is still
producing output.

Two changes, because either alone leaves a hole. WA_QuitOnClose is
cleared on the console, the preferences dialog and the profile
manager, so a log or settings window can never hold the process open
however it came to be visible -- including a user deliberately leaving
one open. And RvApplication::isShuttingDown(), set in ~RvDocument
before the last document closes those windows, gates the auto-show so
no console appears on the way out.

The attribute is applied to all three because they are closed in the
same ~RvDocument block and have exactly the same exposure; fixing only
the one that was observed would leave two identical bugs behind.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
detachAudioOutputDevice() hands work to the audio thread with
Qt::BlockingQueuedConnection three times over -- emitStopDevice(),
emitStopAudio(), and the delete of the output objects -- and none of
them checked that the thread could run it. A BlockingQueuedConnection
blocks until the target's event loop dispatches the call, so a thread
that never started, whose createAudioOutput() failed, or that has
already left exec() never releases the caller. detachAudioOutputDevice()
then never reaches its own quit() and wait(). Calling from the audio
thread itself would deadlock outright.

canBlockOnAudioThread() answers that before each handoff: the thread is
running, it has an event dispatcher, and it is not us. Skipping the
stops costs nothing when it is false, since a thread not running its
loop is not playing either.

The delete becomes best-effort rather than all-or-nothing. It is still
marshalled when there is a loop, which is what keeps Qt6 debug builds
from tripping QObject::~QObject()'s cross-thread assertion on the
QIODevice that QWindowsAudioSink parents. When there is not, the now
idempotent deleteAudioOutputObjects() runs after wait() has returned
and the thread is finished. A Qt warning on the way out beats never
getting out.

wait() is bounded and reports once if the bound is reached, so a wedged
audio thread degrades to a slow exit rather than no exit.

This is not the hang reported on Windows -- a captured stack put that
in RvConsoleWindow -- but it is the same failure waiting to happen, and
it is not reachable by inspection alone.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
front() on an empty container is undefined behaviour, and a hard assert
("front() called on empty vector") in an MSVC debug build. Callers pair
data()/rawData() with size(), so reporting the absence of storage with a
null pointer makes a zero-length copy out of or into a cleared property
a no-op rather than a crash.

Release builds already returned null here for any property that never
held a value, so this only makes the existing behaviour well defined.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
PackageManager deleted m_globalSettingsP without clearing it.
globalSettings() only allocates when that pointer is null, so every
later caller got a reference to freed memory and died dereferencing the
destroyed QSettings inside it.

RvDocument reaches this while closing the last document, so anything
that saves settings after that point -- RvConsoleWindow::done() closing
the console dialog, for one -- crashed on the way out.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
When createAudioOutput() fails the render thread returns without
reaching exec(), so no event loop ever runs on it. detachAudioOutputDevice()
cannot marshal the deletion back onto that thread afterwards -- its
BlockingQueuedConnection has no loop to run on -- so whatever was
allocated before the failure was never freed.

Release it here, on the thread that owns it, before returning.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Two faults in DesktopVideoDevice::queryColorProfile().

The search for a window on the target screen walked every top-level
QWindow in the process, and QWindow::winId() creates the platform window
when there is none. Every QQuickWidget -- so every QWebEngineView panel,
Live Review among them -- owns a parentless offscreen QQuickWindow that
Qt is explicit must never be created ("Do not call create() on
offscreenWindow", qquickwidget.cpp). Handing it a platform window trips
Q_ASSERT(!d->offscreenWindow->handle()) at the end of
QQuickWidget::createFramebufferObject() and aborts RV the moment that
panel is first shown. Consider only windows that are already realized:
any of them on that screen reports the same monitor profile.

The ICC lookup below it sized its path buffer from an unchecked length,
freed a new[] allocation with scalar delete, used UrlCreateFromPath's
output without checking it succeeded, and dereferenced
cmsOpenProfileFromFile's result without a null test -- it returns null
when the path the driver reported is gone or unreadable -- and never
closed the profile. Use std::vector, take the DWORD the API actually
wants, and check both calls.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…indow

The presentation output's ScreenView was a top-level QOpenGLWidget. In
Qt 6 that is composited through its own top-level window's RHI backing
store and takes that window's GL context as its share parent, not the
application's global share context -- so it can land in a private share
group. transfer() wraps the renderer's output FBO colour texture in a
local FBO, and a texture is only visible across contexts in the same
group: in the wrong one glIsTexture() is false for a live texture,
every transfer() is refused and the second display stays black.
Which group it landed in varied run to run, which is what made the
black presentation output intermittent.

Split the class. ScreenWindow is a QOpenGLWindow, which takes the
context to share with as a constructor argument -- the only point at
which sharing can be established. ScreenView becomes a plain QWidget
container around it via createWindowContainer, so the top-level is not
forced onto the OpenGL RHI backend. This mirrors GLView/GLWindow, which
is the main view and demonstrably sits in the renderer's group.
PartialUpdateBlit keeps a backing FBO -- transfer() requires
fboID() to be non-zero -- and does not clear before paintGL().

initializeGL() no longer calls context()->setShareContext(): that only
takes effect on the next create(), and the context already exists by
then, so it never did what it looked like it did. It verifies the share
group and reports a mismatch instead.

Two consequences of the new surface:

- open() makes the share device current before copying its surface
  format. A QOpenGLWindow creates its context lazily, so right after a
  main-view backend swap the new main view's context does not exist yet
  and glShareContext() is null -- which is another way to end up outside
  the renderer's group.
- Nothing primes the context in open() any more. A PartialUpdateBlit
  window's makeCurrent() binds the backing FBO that Qt only creates on
  the first paint, so calling it before then dereferences a null FBO
  inside Qt and takes the process down. transfer(), transfer2() and
  makeCurrent() gate on fboID() for the same reason and skip the frame;
  show() drives the expose that creates the FBO.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
FBOs are not shared between contexts but textures are, so transfer()
and transfer2() keep a per-context clone of the renderer's output FBO
wrapped around its colour texture. The cache was keyed on the source
pointer and trusted forever. Two ways that goes wrong:

- The renderer deletes and reallocates those FBOs
  (ImageRenderer::Device::clearFBOs, ImageFBOManager::newImageFBO), so
  the same address comes back as a different FBO.
- The borrowed texture can be dead by the time it is attached, which
  leaves the clone incomplete without any call failing outright.

Nothing ever invalidated a cached incomplete clone, so a single
transient error turned into a permanently black output that every later
frame blitted from.

Route both paths through cloneForSource(), which re-verifies a cached
clone against the source's texture, target and size, discards an
incomplete clone rather than caching it, and returns null so the caller
skips the frame. On the first refusal it reports whether the texture is
gone or merely in another share group, which are otherwise
indistinguishable from the outside.

GLFBO::isComplete() is the non-throwing completeness test that needs,
for callers assembling an FBO from attachments they do not own.
releaseFBOClones() drops the cache with a context current, and
clearCaches()/unbind() now go through it. TWK_GLDEBUG after
glBlitFramebuffer so an incomplete framebuffer is attributed to the
blit instead of to the next frame's makeCurrent().

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Quitting out of presentation mode logged a long tail of
GL_INVALID_OPERATION starting in ~GLFBO, and left the presentation
output black for the rest of the session. All of it is GL teardown
running with no context current: once none is, glGetError() keeps
returning that same error, so one lost context is worth a great many
messages, none of them near the cause.

- QTGLVideoDevice::makeCurrent() did nothing at all when the platform
  surface was gone but the QOpenGLContext was not -- which is the state
  Qt leaves the device in while shutting down, since the native window
  is destroyed before the C++ object. Keep a QOffscreenSurface, created
  while the window is still healthy, and bind the context to that
  instead: GL deletion needs a current context, not a visible one. The
  handle() test also belongs in the outer condition rather than nested
  inside it, where a live window with a dead surface fell through every
  branch silently. The genuinely unreachable case now says so once.

- DesktopVideoDevice::close() deleted the view before m_viewDevice.
  ~GLVideoDevice deletes the device's GL text context, and
  ~GLTextContext deletes the FTGL fonts, which delete GL textures --
  all of it against a context the view had just taken with it, so the
  textures leaked on every presentation-mode toggle, not only at exit.
  Release the FBO clones first (that needs the view's context), then
  the device, then the view, then hand the main view's context back.

- VulkanDesktopVideoDevice::close() deliberately does not chain to the
  base, so it has to release the FBO clones itself -- before
  setViewDevice(nullptr) takes away the device it needs to make a
  context current.

- RvApplication dropped the session's output device only after
  close()ing it. Unbind first: ImageRenderer::setOutputDevice() calls
  unbind() on the outgoing device, and it has to run while that
  context is alive. It then restores the main view's context, because
  close() leaves nothing current and DesktopVideoDevice's own restore
  goes through its share device, which is null whenever the main view
  is Vulkan. That is where the bogus "Could not retrieve OpenGL
  version. Make sure you have installed the Nvidia drivers." came
  from: queryGLIntoContainer() reading GL_VERSION with no context.

- ImageRenderer::setOutputDevice() falls back to the control device's
  context. m_outputDevice.glDevice is a dynamic_cast to GLVideoDevice
  and is null for every GLBindableVideoDevice output -- presentation,
  AJA, NDI -- because the two are siblings, not base and derived. The
  control context is the right fallback: it owns the FBOs being cleared,
  and the rest of the function already depends on it further down.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
DesktopVideoModule::rebuildDevices() re-derived the GL-vs-Vulkan
decision from shouldUseVulkanPresentation(), which reads the persisted
display-depth preference. That is the requested intent, not the backend
the main view is actually running: a 10-bit request that fell back to
GL at runtime keeps its 10-bit intent on purpose. A presentation output
built on the opposite backend to the viewport is a black second
display.

Pass the backend down from RvDocument instead. It is the only place
that knows which widget actually exists now, and it calls
rebuildDesktopVideoDevices() from each of the three swap paths anyway.
shouldUseVulkanPresentation() stays for the initial build, when there
is no main view to ask, and now says so.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
RvDocument's "OpenGL is already live" branch wrote neither Options nor
QSettings, and returns early whenever the GL context already has the
requested depth. The 10-bit and Vulkan-live branches above it both
write those first.

Selecting 8-bit from 10-bit therefore did nothing to the very state
other subsystems read back as "the requested display depth", leaving
Options claiming 10-bit for the rest of the session. Write the depth
before anything below can return early.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Three sizing faults, all of which surface once the output lives on a
screen whose devicePixelRatio differs from the one its window was
created on.

- The swapchain-recreate test compared the caller's requested size
  against m_vkSwapchainExtent. createSwapchain() takes its extent from
  capabilities.currentExtent, so it cannot be driven to match a request
  the surface disagrees with: any caller off by even a pixel recreated
  the swapchain on every frame, forever and silently, since both
  surface-format reports are latched. Compare against the surface's
  current extent instead.

- The grow-only shared image sized its headroom from the primary
  screen. Use this window's own screen: a presentation output lives on
  a second display, and the primary is both the wrong one and, when it
  is the smaller of the two, useless as headroom.

- The present copy and blit took their destination extent from the
  shared image, which is sized from the caller's request. A stale
  devicePixelRatio inflating that -- a 3840x2160 output asking for
  5760x3240 -- wrote outside the swapchain image, which is invalid
  usage and so undefined contents or a faulted submit rather than a
  visible error. vkCmdCopyImage cannot scale, so clamp it to the
  overlap; vkCmdBlitImage can, so fill the swapchain from the used
  sub-region.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Fixes from a C++ review of this branch.

- twkGlAnyContextIsCurrent() probed with glGetString(), which is
  undefined with no context current and can crash on macOS. Ask the
  platform instead (CGLGetCurrentContext / wglGetCurrentContext); Linux
  keeps glGetString(), which GLVND answers with null.

- When the audio thread does not exit within the bounded wait, stop
  deleting objects it still owns and leak the thread, detached from its
  parent, instead of destroying a running QThread (fatal in Qt6).

- presentPixelData() compared the swapchain against the requested size
  rather than the surface extent, recreating it every frame on a
  mismatch and copying past the swapchain image. Share the surface
  check with getSharedImageInfo() and clamp the copy.

- Check every staging buffer create/allocate/bind/map result and null
  freed handles, so a failed map no longer writes through an
  uninitialised pointer and an early return cannot double free.

- A failed submit after the fence reset left the fence unsignaled, so
  the next wait on that slot hung forever. Re-signal it with an empty
  submit that also consumes the acquire semaphore, then fall back to GL.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Follow-up to the C++ review of this branch.

- Order interop tiling candidates by vendor preference: OPTIMAL first
  on NVIDIA, LINEAR first elsewhere. Mesa reports OPTIMAL as exportable
  but renders tile garbage with it, so trying it first everywhere
  regressed RADV. Drop the now redundant legacy-heuristic log line.

- ~QTVulkanVideoDevice cleans every slot whenever its context exists;
  slot 1 imports could previously leak and pin the Vulkan memory.

- Document that GLContextScope keeps an already-current context, which
  is right for shared objects but not for FBOs.

- Store exported fds/handles in the slot record as soon as they are
  obtained, so a later failure in getSharedImageInfo() closes them.

- Delete copy operations on QTVulkanVideoDevice and QTGLVideoDevice,
  and close an existing view in VulkanDesktopVideoDevice::open().

- rebuildDesktopVideoDevices() takes the initiating document's session
  instead of the active document's.

- Extract DesktopVideoDevice::tenBitDisplayRequested() and persist the
  display depth once in setDisplayOutput().

- Make the report-once flags atomic, guard a negative screen index,
  catch by const reference, fix narrowing and -Wparentheses, and
  initialise the new RvDocument members in declaration order.

- Remove dead code (unused instance setup, duplicated kMaxStaleSeconds,
  VulkanBuildProbe.cpp), give findMemoryType internal linkage with an
  unsigned shift, and use include guards in the new RvCommon headers.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
The fourth parameter of UrlCreateFromPath is a DWORD dwFlags (reserved,
must be 0), not a pointer. The NULL -> nullptr sweep in b750556 turned
the original NULL (which MSVC defines as 0) into nullptr, which does not
convert to DWORD and fails the Windows build with C2664.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…k into VulkanWindow

- VulkanWindow/VulkanView/RvDocument/VulkanDesktopVideoDevice: move
  members not set from ctor params to default member initializers.
- VulkanView owns QTVulkanVideoDevice via std::unique_ptr, still reset
  explicitly at the same point in the destructor.
- QTVulkanVideoDevice: GL context, offscreen surface, FBO and translator
  held by std::unique_ptr with the explicit release order kept;
  ~QTVulkanVideoDevice() override.
- Group per-frame parallel arrays into FrameSync, StagingBuffer and
  SharedImage structs (VulkanWindow) and SharedGLObjects
  (QTVulkanVideoDevice).
- File-static helpers into an anonymous namespace; findMemoryType
  returns std::optional; deviceProc<> helper for vkGetDeviceProcAddr;
  transitionImageLayout() helper for the image barriers.
- std::numeric_limits instead of UINT32_MAX/UINT64_MAX; std::array for
  candidates, composite alpha preference and submit/present arrays;
  std::string_view name helpers; descriptive local names.
- using instead of typedef for Timer; range-based for where the index
  is unused.
- Remove unused RvDocument::vulkanView(), VulkanView::vulkanWindow(),
  VulkanView::isInitialized(), QTVulkanVideoDevice::vulkanWindow() and
  QTVulkanVideoDevice::eventWidget().
- Remove ImageRenderer::reportGL(), replaced by debugGpu().

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
…nWindow

- forcedTilingRequested() returns std::optional<VkImageTiling> instead of
  a bool plus an out-parameter
- QTVulkanVideoDevice::setEventWidget uses the same ternary as the
  constructor

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Present the main view, and the Presentation Mode second-display output,
through Metal on macOS (IOSurface + CALayer) when a 10-bit display depth is
requested: the macOS counterpart of the Vulkan backend on Linux and Windows.

- MetalView / QTMetalVideoDevice: a native CALayer-backed QWidget whose
  IOSurface is rendered by the GL-on-Metal context through IOSurface GL
  interop, with a GPU-blit CPU fallback (RV_METAL_FORCE_CPU_PRESENT=1 forces
  it for testing). Set RV_METAL_DEBUG_PRESENT to trace presents.
- MetalDesktopVideoDevice: the 10-bit presentation output, a passive
  fullscreen MetalView whose device is the base m_viewDevice, so the
  inherited transfer()/transfer2() and every stereo mode work unchanged.

Shares one backend-selection path with Vulkan rather than duplicating it:

- DesktopVideoDevice::shouldUseNativePresentation() (was
  shouldUseVulkanPresentation) picks Vulkan or Metal per platform, for both
  the main view and the presentation output.
- RvDocument routes Metal through the backend-neutral helpers
  (viewVideoDevice(), setActiveViewContentSize(), ...) and adds
  swapGLViewToMetal() / fallbackMetalToGLView() beside the Vulkan swaps, so
  a display-depth change applies at runtime. setDisplayOutput() dispatches
  to whichever native backend the platform has.
- The desktop share device is retyped to TwkGLF::GLVideoDevice so the Metal
  main view can be it; the Vulkan path still passes null.
- DesktopVideoModule::rebuildDevices() now retires the replaced devices and
  RvApplication::rebuildDesktopVideoDevices() purges them only after
  IPGraph::refreshPhysicalDevices() has re-pointed the display groups and the
  control device's physical device. Closing a device destroys a native
  window, which pumps the event loop and repaints through the graph, so
  destroying them first was a use-after-free.
- fix(ipcore): setPhysicalDevicesInternal() deleted display groups by index
  while each destructor erased itself from the same vector, skipping every
  other group and leaving dangling device pointers. Iterate over a copy.

Squashed from the metal-10bit-macos development history and rebased onto
the Vulkan presentation branch.

Signed-off-by: Cédrik Fuoco <cedrik.fuoco@autodesk.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant