
Description
Run and interpret the complete read-only dgx-assist diagnostic suite for NVIDIA DGX Station GB300, correlate findings with pinned NVIDIA playbooks, export a redacted support bundle, and apply one separately approved allowlisted fix. Use when the user reports a Station, CUDA, GPU health, coherency, vsloshd, Docker, CDI, MIG, cache, port, or owned inference-service failure.
SKILL.md
DGX Station diagnostics
Diagnose first. Do not mutate as part of diagnosis.
Workflow
- Run
scripts/dgx-assist diagnose run --json. - Report the detected compatibility profile, then findings in severity order with their stable IDs and evidence. Preserve
unknownstates. On Software 1.0, do not reinterpret intentionally skipped Software 2.0 service checks as faults. - Search the pinned playbooks for each high or critical finding; cite the relevant URL, heading, lines, and commit.
- If no passage overlaps, say so and avoid inventing a platform fix.
- Offer
diagnose bundle --report-id "<id>"when escalation is appropriate. - Offer at most one automatic fix at a time, and only when
fix_idis present. - Preview with
diagnose fix --report-id "<id>" --finding "<id>" --dry-run. - Explain exact actions, impact, privilege, and reboot state. Obtain explicit approval.
- Repeat with
--yesonly after approval and report the action receipt.
Safety requirements
- Keep
diagnose runread-only. - Never install packages, rewrite Docker configuration, change power caps or driver parameters, kill workloads, or modify MIG through a diagnostic fix.
- Never stop a service without current
dgx-assistownership evidence. - Never use Fabric Manager as a routine Station check or remediation.
- Re-run diagnostics when finding evidence is stale.
- Never reveal secrets or unredacted home paths in a bundle.
- Do not execute an unregistered remediation.
- Treat
--yesonly as approval already obtained.
Read references/findings.md before proposing a fix or support bundle. Read references/bringup.md when the problem concerns physical deployment, BMC or firmware verification, driver bring-up, power braking, or support escalation.
More skills from the dgx-spark-playbooks repository
View all 15 skillsanalysis-methods
write Python analysis code for FHIR data
Jul 14Data AnalysisFHIRHealthcareNVIDIA +1case-summary
summarize clinical patient cases from FHIR
Jul 14FHIRHealthcareNVIDIASummarizationclinical-delegation
delegate clinical tasks to specialist agents
Jul 14AgentsHealthcareMulti-AgentNVIDIAclinical-knowledge
provide clinical reference and regulatory context
Jul 14Clinical TrialsHealthcareNVIDIARegulatory Compliancecohort-compare
analyze patient cohorts from FHIR endpoints
Jul 14Data AnalysisFHIRHealthcareNVIDIAdgx-diagnose
diagnose NVIDIA DGX Station hardware issues
Jul 14AI InfrastructureDebuggingNVIDIAObservability
More from NVIDIA
View publishernemoclaw-user-guide
retrieve NemoClaw documentation and configuration
NemoClaw
Jul 20DocumentationMCPSearchmcore-build-and-dependency
manage Megatron-LM development environments
Megatron-LM
Jul 27ContainersDeploymentPythonmcore-bump-base-image
update NVIDIA PyTorch base images
Megatron-LM
Jul 14CI/CDDeploymentmcore-cicd
manage CI/CD pipelines for Megatron-LM
Megatron-LM
Jul 27CI/CDDeploymentGitHubmcore-create-issue
investigate CI failures and create issues
Megatron-LM
Jul 14DebuggingGitHubTriagemcore-linting-and-formatting
lint and format Megatron-LM code
Megatron-LM
Jul 14Best PracticesCode Analysis