Skip to content

{CUDA Decoding} Restore SW decoder fallback in xprsDecoderMaker - #262

Open
YLouWashU wants to merge 1 commit into
facebookresearch:mainfrom
YLouWashU:export-D114806413
Open

{CUDA Decoding} Restore SW decoder fallback in xprsDecoderMaker#262
YLouWashU wants to merge 1 commit into
facebookresearch:mainfrom
YLouWashU:export-D114806413

Conversation

@YLouWashU

@YLouWashU YLouWashU commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Summary:
A wheel built with ENABLE_NVCODEC=ON segfaults on any machine without an NVIDIA driver. The GitHub CI "Test wheel" job crashes with exit 139 after:

Cannot load libcuda.so.1
[XPRS][WARNING]: Loading CUDA functions failed
[DecoderFactory][WARNING]: Could not create a decoder for 'H.265'!

With no driver, getNvCodecContext() throws while probing the HW entries and enumDecoders() catches it, returning a non-OK result. The SW hevc/h264 decoders were already collected before the throw (they precede the HW entries in kPreferredDecoderImplementations), so the list is usable. xprsDecoderMaker bailed on the non-OK result and discarded that list, so no decoder was created and the null decoder was dereferenced downstream. Note the HW decoders never enter the list at all: the throw happens before codecs.push_back.

Three changes:

  • Ignore the enumDecoders() result and fail only when the decoder list is genuinely empty. This is what fixes the crash.
  • On decoder init() failure, continue to the next candidate instead of returning, so the HW->SW fallback actually runs. This matters on a machine that does have a GPU, where enumeration succeeds but init() fails (NVDEC session limit, VRAM pressure).
  • Return nullptr rather than decoder when the loop is exhausted. Raised by DevMate: with the continue above, a candidate that initialized but failed decode() could otherwise be returned by a later iteration's fall-through. A non-null return is also load-bearing in DecoderFactory::makeDecoder, which stops at the first maker that returns anything, so handing back a known-bad decoder would prevent other registered makers from being tried. This also removes a pre-existing path to the same stale return.

The GPU path is unchanged.

Reviewed By: kongchen1992

Differential Revision: D114806413

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 4, 2026
@meta-codesync

meta-codesync Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

@YLouWashU has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114806413.

@YLouWashU
YLouWashU force-pushed the export-D114806413 branch 2 times, most recently from 195900f to be8d7d9 Compare August 4, 2026 23:56
Summary:
A wheel built with `ENABLE_NVCODEC=ON` segfaults on any machine without an NVIDIA driver. The GitHub CI "Test wheel" job crashes with exit 139 after:

    Cannot load libcuda.so.1
    [XPRS][WARNING]: Loading CUDA functions failed
    [DecoderFactory][WARNING]: Could not create a decoder for 'H.265'!

With no driver, `getNvCodecContext()` throws while probing the HW entries and `enumDecoders()` catches it, returning a non-OK result. The SW `hevc`/`h264` decoders were already collected before the throw (they precede the HW entries in `kPreferredDecoderImplementations`), so the list is usable. `xprsDecoderMaker` bailed on the non-OK result and discarded that list, so no decoder was created and the null decoder was dereferenced downstream. Note the HW decoders never enter the list at all: the throw happens before `codecs.push_back`.

Three changes:
- Ignore the `enumDecoders()` result and fail only when the decoder list is genuinely empty. This is what fixes the crash.
- On decoder `init()` failure, continue to the next candidate instead of returning, so the HW->SW fallback actually runs. This matters on a machine that does have a GPU, where enumeration succeeds but `init()` fails (NVDEC session limit, VRAM pressure).
- Return `nullptr` rather than `decoder` when the loop is exhausted. Raised by DevMate: with the `continue` above, a candidate that initialized but failed `decode()` could otherwise be returned by a later iteration's fall-through. A non-null return is also load-bearing in `DecoderFactory::makeDecoder`, which stops at the first maker that returns anything, so handing back a known-bad decoder would prevent other registered makers from being tried. This also removes a pre-existing path to the same stale return.

The GPU path is unchanged.

Reviewed By: kongchen1992

Differential Revision: D114806413
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant