calesthio/nvidia-maxine-audio-effects
Use NVIDIA Maxine / NVIDIA AI for Media audio effects for speech cleanup and enhancement in live or offline media pipelines, including Background Noise Removal, denoise+dereverb, room echo removal, acoustic echo cancellation, audio super-resolution, Studio Voice, Speaker Focus, and Voice Font. Use when selecting, integrating, prompting around, QAing, or privacy-reviewing NVIDIA AFX SDK or BNR NIM workflows for conferencing, broadcast, podcast, avatar, ad, social, and video post-production audio.
npx skills add https://github.com/calesthio/generative-media-skills --skill nvidia-maxine-audio-effects
Use this skill to plan and operate NVIDIA Maxine Audio Effects, now under NVIDIA AI for Media, as an audio cleanup and enhancement stage. Treat it as a production and integration skill, not as a sound-generation prompt recipe: the core work is choosing the correct effect, preparing audio correctly, protecting speaker rights, running the SDK or NIM in the right runtime, and proving that the processed speech is clearer without damaging the performance.
Facts below were verified against official NVIDIA documentation on 2026-07-10. Re-check NVIDIA docs, NGC, and license terms before quoting exact model profiles, supported GPUs, SDK versions, file limits, or access requirements in a production proposal.
NVIDIA documents these Audio Effects SDK capabilities:
denoiser / Background Noise Removal (BNR): remove common background noises from speech while preserving natural speech as much as possible.dereverb: suppress room echo/reverberation from recordings made in reflective rooms.dereverb_denoiser: use the combined denoise + dereverb model when both noise and reverb are present; do not chain separate denoiser and dereverb effects when NVIDIA provides the combined effect for that case.aec: acoustic echo cancellation for live bidirectional communication. This is not the same as dereverb; it needs near-end microphone audio plus a far-end/reference signal.superres: audio super-resolution, typically 8 kHz -> 16 kHz or 16 kHz -> 48 kHz, to restore/predict missing high-frequency content.studio_voice_high_quality: offline enhancement for degraded speech from poor microphones, noise filters, beamforming, static, or non-ideal acoustics.studio_voice_low_latency: real-time Studio Voice variant; use for live conferences or broadcasts, not offline batch finishing.speaker_focus: Early Access speaker isolation for keeping the primary speaker while suppressing other speakers and some noises.voice_font_high_quality / voice_font_low_latency: Early Access any-to-any voice conversion using reference speech. Treat as voice identity transformation requiring explicit consent and review.NVIDIA BNR NIM is narrower than the full AFX SDK. Use it for Background Noise Removal through a containerized gRPC service or hosted Try API; do not assume NIM exposes dereverb, AEC, super-resolution, Studio Voice, Speaker Focus, or Voice Font unless current NVIDIA docs say so.
Do not use Maxine audio effects when the user needs:
Select the deployment path before designing the cleanup chain.
Use the AFX SDK when the application needs one or more SDK effects beyond BNR, local custody, low-latency client integration, or direct control of sample format, batching, and chained effects. NVIDIA positions Windows SDK packages for client-side applications and Linux SDK packages for server/datacenter/cloud deployments. The SDK requires NVIDIA GPUs with Tensor Cores; official docs list Turing, Ampere, Ada, Hopper, and Blackwell families for current Linux SDK support.
Operational implications:
Use BNR NIM when the job is specifically Background Noise Removal and the team wants a containerized service with gRPC, cloud/datacenter deployment, health checks, observability, and scalable inference.
BNR NIM has two modes:
Operational implications:
NGC_API_KEY, and a compatible NIM_MODEL_PROFILE when selecting a profile explicitly.Start from the artifact, not the tool name.
| Input problem | Recommended Maxine path | Avoid |
|---|---|---|
| Fan, keyboard, traffic, HVAC, household noise under speech | BNR / denoiser; consider BNR 2.0 for ASR-prep if current docs and access support it | Heavy intensity before checking consonants, breaths, and emotion |
| Roomy lavalier, echoey webinar, reflective office | dereverb; if noise is also present use dereverb_denoiser | Using AEC for room reverb; AEC solves far-end playback echo |
| Remote guest hears their own delayed voice in a live call | aec with near-end mic plus far-end reference signal | Offline dereverb-only cleanup |
| Narrowband phone/Zoom archive that must become 48 kHz delivery audio | superres after basic cleanup, or a supported chain such as superres -> denoise where documented | Claiming real high-frequency recovery; describe it as inferred enhancement |
| Cheap headset or laptop mic sounds thin/static/processed | Studio Voice High Quality for offline; Low Latency for live | Low-latency Studio Voice for batch post-production |
| Main presenter has other people talking in background | speaker_focus only if Early Access/access constraints are acceptable and the clip has a clear primary speaker | Promising perfect diarization or multi-speaker preservation |
| User asks to make one person sound like another | Voice Font only with explicit rights/consent, reference custody, and disclosure policy | Covert voice conversion, impersonation, or using unlicensed reference speech |
For podcast, ad, avatar, or social pipelines, make Maxine one stage in a broader audio finish:
Documented effect constraints override these heuristics.
dereverb_denoiser over separate denoise and dereverb.Before running an effect:
Use the least processing that solves the content problem. NVIDIA exposes an intensity_ratio for several workflows where 0 is effectively passthrough and 1 is maximum impact. In production, render a small A/B ladder before committing:
0.35-0.50: light cleanup for premium voiceover, interview, or emotional performance.0.60-0.80: typical remote guest, webinar, or social clip cleanup.0.85-1.00: rescue audio where intelligibility matters more than naturalness.Listen for over-processing:
If artifacts appear, reduce intensity, split the clip into regions, bypass clean sections, try native sample-rate cleanup before super-resolution, or use conventional editing for clicks/clips instead of stronger AI cleanup.
Use low-latency paths only. For a two-way call, route microphone audio through BNR/dereverb/Studio Voice Low Latency as needed and route echo cancellation through AEC with the far-end signal. Measure round-trip latency; do not add high-quality offline Studio Voice or large-buffer batch steps to a live path.
Process each participant independently where possible. Use BNR for background noise, dereverb_denoiser for remote rooms with both hiss and reverb, and Speaker Focus only when background speakers are disposable. Preserve laughter, sighs, and emotional tone unless the brief says to sanitize them. Normalize loudness only after AI cleanup.
For recorded human VO, prioritize naturalness over maximum suppression. Use light-to-medium BNR or Studio Voice High Quality, then EQ/compress. For synthetic/avatar voice tracks, avoid heavy denoise unless there is actual generated noise; many artifacts are vocoder issues better fixed upstream by regenerating the voice.
Use stronger cleanup if the platform and content tolerate a processed sound. Create a 10-20 second proof segment first. Preserve viral context noises if they matter to the joke/story; do not remove everything that is not speech by default.
When audio is being cleaned before transcription, prioritize word error reduction over studio tone. BNR 2.0 is documented as a newer experimental denoiser focused on ASR accuracy; use it only if the current SDK/NIM path supports it and label the choice experimental. Always run before/after ASR spot checks on names, numbers, and code-switching.
Audio often contains biometric voice data, private conversations, unreleased scripts, and background speech from bystanders. Before processing:
For every Maxine processing run, record:
Pass a processed file only when it satisfies all applicable checks:
Use both objective and subjective checks. Objective checks can include waveform duration, LUFS, peak/true peak, spectrogram inspection, ASR before/after on a representative segment, and dropped-frame/sync checks after video mux. Subjective checks require headphones and at least one loudspeaker pass because denoise artifacts often reveal themselves differently.
Production intent: clean a 42-minute remote podcast guest recorded in a kitchen with fridge hum, laptop fan, and room reverb.
Approach:
dereverb_denoiser at intensity 0.55, 0.70, and 0.85.Why: the combined effect is designed for simultaneous noise and reverb; a proof ladder avoids over-processing a long performance.
Expected result: clearer speech and reduced room tone while retaining the guest's personality.
Failure modes: watery vowels at high intensity, clipped plosives that no denoiser can fix, background clatter that survives the first seconds, and ASR regressions on proper nouns.
Production intent: improve a live panel webinar without adding noticeable delay.
Approach:
Why: AEC solves far-end echo, while BNR/dereverb clean the microphone path. Offline/high-quality models are not appropriate for live latency budgets.
Expected result: intelligible panel audio with less feedback/noise and no major delay.
Failure modes: AEC without a correct far-end reference, over-suppression during double-talk, and latency spikes when concurrency exceeds GPU resources.
Production intent: batch-clean short social clips before captioning, using a self-hosted BNR NIM on an internal GPU server.
Plan:
NGC_API_KEY, selected compatible NIM_MODEL_PROFILE, published service port, and TLS/mTLS if clients are remote.Why: NIM gives service deployment and observability while keeping audio inside the organization's infrastructure.
Expected result: cleaner captions and fewer ASR errors without uploading unreleased social assets to a hosted preview endpoint.
Failure modes: wrong model profile for the GPU, WAV/sample-rate mismatch, unsecured internal service, or treating preview API limits as production guarantees.
Take calesthio/nvidia-maxine-audio-effects from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.