nvidia/flashdreams-postprocessing
Add or modify FlashDreams video post-processing processors, sessions, presets, and runner stream wiring. Use when implementing a new VideoPostProcessorConfig / VideoPostProcessor / VideoPostProcessorSession, registering a --postprocess.preset entry point, changing VideoPostprocessStream behavior, or reasoning about streaming buffering, layouts, per-view processing, distributed execution, or postprocess tests.
npx skills add https://github.com/NVIDIA/flashdreams --skill flashdreams-postprocessing
Use this skill when adding a video post-processor or changing the runner
post-processing stream. The reference implementation is
integrations/flashvsr/flashvsr/postprocess.py.
A post-processor is usually three classes, not one class inheriting everything:
VideoPostProcessorConfig: serializable config and CLI surface. It sets_target to the processor factory and declares fields, output_spec(),
requires_all_ranks(), and validate_execution().
VideoPostProcessor: lightweight factory created from config. Its job isstart(spec) -> VideoPostProcessorSession.
VideoPostProcessorSession: mutable per-stream runtime. It owns buffers,caches, lazy model instances, counters, and process() / flush().
Keep stream state in the session. Do not store per-rollout mutable state on the
config or processor factory.
flashdreams/flashdreams/infra/postprocess/.integrations/<name>/<pkg>/postprocess.py.
@dataclass(kw_only=True)
class MyPostProcessorConfig(VideoPostProcessorConfig):
_target: type["MyPostProcessor"] = field(
default_factory=lambda: MyPostProcessor
)
scale: int = 2
def output_spec(self, input_spec: VideoSpec) -> VideoSpec:
return VideoSpec(
height=input_spec.height * self.scale,
width=input_spec.width * self.scale,
fps=input_spec.fps,
channels=input_spec.channels,
)
Override:
output_spec() when spatial size, channels, or timing changes.requires_all_ranks() when the processor must run on nonzero ranks undertorchrun.
validate_execution() to reject unsupported distributed or shape modesearly.
class MyPostProcessor(VideoPostProcessor[MyPostProcessorConfig]):
def start(self, spec: VideoSpec) -> VideoPostProcessorSession:
return _MyPostProcessorSession(self.config, spec)
class _MyPostProcessorSession(VideoPostProcessorSession):
def __init__(self, config: MyPostProcessorConfig, spec: VideoSpec) -> None:
self._config = config
self._spec = spec
self._buffer: Tensor | None = None
def process(self, chunk: VideoChunk) -> list[VideoChunk]:
...
def flush(self) -> list[VideoChunk]:
...
process() is synchronous but may return []: that means it consumed the
input chunk and is buffering frames until a later chunk or flush() can
complete an output window.
VideoChunk.tensor in chunk.layout.to_bvtchw() only as a generic boundary helper.model kernels require that, and keep internal buffers in that native
layout.
.contiguous() because it can copy.VideoChunks:[-1, 1] value range unless the API is intentionally changed.layout. [project.entry-points."flashdreams.postprocess_presets"]
"my-postprocessor-v1" = "my_pkg.postprocess:POSTPROCESS_PRESET_MY_V1"
The exported object must be a VideoPostProcessorConfig, for example:
POSTPROCESS_PRESET_MY_V1 = MyPostProcessorConfig(...)
Users select it with --postprocess.preset my-postprocessor-v1.
Runners create a VideoPostprocessStream through
create_runner_postprocess_stream(). The stream:
view when postprocess_per_view=True;
session.process(VideoChunk(...)) for each AR output;[] into a zero-frame tensor so process() remains tensor-only;_append_if_nonempty();flush() once at end-of-stream and appends any tail output.Use postprocess_output_layout to describe the runner's decoded output layout.
Use postprocess_per_view=True for bvtchw outputs when each camera/view needs
an independent processor session.
Add CPU-safe tests unless the behavior genuinely requires a GPU:
flashdreams/tests/test_postprocess_presets.py.flashdreams/tests/test_postprocess_stream.py.integrations/<name>/tests/test_postprocess.py.flashdreams/tests/test_runner_postprocess.py.
Every pytest test must use exactly one marker: ci_cpu, ci_gpu, or manual.
Prefer fake processor builders for CPU tests instead of loading checkpoints.
Useful focused validation:
uv run pytest flashdreams/tests/test_runner_postprocess.py \
flashdreams/tests/test_postprocess_stream.py \
flashdreams/tests/test_postprocess_presets.py \
integrations/<name>/tests/test_postprocess.py
Take nvidia/flashdreams-postprocessing from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.