[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"skill-aws-labs-hyperpod-incident-triage":3,"mdc--i2xopr-key":48,"related-org-aws-labs-hyperpod-incident-triage":447,"related-repo-aws-labs-hyperpod-incident-triage":626},{"slug":4,"name":4,"fn":5,"description":6,"org":7,"tags":12,"stars":26,"repoUrl":27,"updatedAt":28,"license":29,"forks":30,"topics":31,"repo":43,"sourceUrl":46,"mdContent":47},"hyperpod-incident-triage","triage SageMaker HyperPod incidents","Correlation and skip rules for SageMaker HyperPod incident triage. Keeps distinct fault types on the same instance group as separate investigations (the default correlator merges them), and prevents periodic-audit re-investigation of an unchanged cluster. Applies at the Incident Triage stage before any investigation runs.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},"aws-labs","AWS Labs","https:\u002F\u002Fpexgzepcugksgbtrxkhf.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Forg-logos\u002Faws-labs.png","awslabs",[13,17,20,23],{"name":14,"slug":15,"type":16},"Operations","operations","tag",{"name":18,"slug":19,"type":16},"Incident Response","incident-response",{"name":21,"slug":22,"type":16},"Triage","triage",{"name":24,"slug":25,"type":16},"AWS","aws",472,"https:\u002F\u002Fgithub.com\u002Fawslabs\u002Fawsome-distributed-ai","2026-09-02T07:47:39.537858",null,206,[25,32,33,34,35,36,37,38,39,40,41,42],"distributed-inference","distributed-training","efa","eks","generative-ai","gpu","hyperpod","kubernetes","parallelcluster","physical-ai","slurm",{"repoUrl":27,"stars":26,"forks":30,"topics":44,"description":45},[25,32,33,34,35,36,37,38,39,40,41,42],"Collection of best practices, reference architectures, model training examples and utilities to train large models on AWS. ","https:\u002F\u002Fgithub.com\u002Fawslabs\u002Fawsome-distributed-ai\u002Ftree\u002FHEAD\u002Farchitectures\u002Fsagemaker-hyperpod-slurm\u002Ftools\u002Fdevops-agent\u002Fskills\u002Fhyperpod-incident-triage","---\nname: hyperpod-incident-triage\ndescription: Correlation and skip rules for SageMaker HyperPod incident triage. Keeps distinct fault types on the same instance group as separate investigations (the default correlator merges them), and prevents periodic-audit re-investigation of an unchanged cluster. Applies at the Incident Triage stage before any investigation runs.\nmetadata:\n  version: \"0.7.0\"\n  agent_types: [\"INCIDENT_TRIAGE\"]\n---\n\n# HyperPod incident triage\n\nUse these rules when deciding whether a HyperPod incident should be **linked**\nto an existing investigation, **skipped**, or investigated on its own. They\nrefine the default correlation for HyperPod's fault model. When a rule below\ndoes not clearly apply, fall back to your normal triage judgment.\n\nEvery HyperPod incident carries a description built by the webhook bridge. Two\nkinds arrive:\n\n- **Cluster events** (a real HyperPod fault) — the description contains the\n  instance group, a `Description:`, and often a `FailureMessage:`.\n- **Periodic audits** (a scheduled health sweep) — the description begins with\n  either \"Periodic audit for HyperPod cluster '\u003Cname>' detected N issue(s)\"\n  (issues found) or \"Daily heartbeat audit for HyperPod cluster\" (the once-daily\n  all-clear liveness signal). Both ask the RCA skill to run in audit mode.\n\n## Linking rules for HyperPod fault events\n\nLink a new HyperPod fault only to an existing HyperPod investigation that\ndescribes **the same fault on the same component**. Treat the instance group\nplus the fault text (`Description:` together with any `FailureMessage:`) as the\nidentity of the fault.\n\n- **Do link** when both the instance group and the fault text match an open\n  investigation — it is the same problem arriving again.\n- **Do NOT link** when the instance group is the same but the fault text\n  differs. Different fault types on the same instance group have different root\n  causes and need separate investigations. For example, a lifecycle-script\n  failure on `worker2` is **not** the same incident as a GPU\u002FXid fault on\n  `worker2` — keep them separate even though they share the instance group.\n- **Do NOT link** the same fault type across different instance groups (e.g.\n  an Xid fault on `worker2` and an Xid fault on `worker3` are two distinct\n  hardware faults).\n- **Do NOT link** capacity errors for different instance types (e.g.\n  `ml.g5.8xlarge` vs `ml.p5.48xlarge`) — the `FailureMessage` differs, so they\n  are distinct problems.\n\nWhen in doubt, prefer a separate investigation over an incorrect link:\nover-linking hides distinct problems, which is the failure this skill exists to\nprevent.\n\n## Rules for periodic audits\n\nA periodic audit is a scheduled sweep, not a specific fault. Decide based on\nwhether the cluster's condition has **changed** since the last audit:\n\n- **Skip** the **daily heartbeat \u002F all-clear audit** (description begins \"Daily\n  heartbeat audit for HyperPod cluster\", or otherwise reports no open issues) —\n  it carries no fault to investigate; the task's presence in the console is\n  itself the liveness signal.\n- **Skip** the audit if another periodic-audit investigation for the same\n  cluster is **currently in progress** — let it finish rather than starting or\n  re-activating a parallel one.\n- **Link** the audit to the most recent **completed** periodic-audit\n  investigation only if the cluster's open problems are **unchanged** — the\n  same fault chains, the same CrashLoopBackOff pods, and the same NotReady\n  nodes as that prior audit found. There is nothing new to report.\n- **Proceed** (investigate) if the cluster shows **any new or changed problem**\n  since the last audit — a new fault event, a newly CrashLooping pod, or a\n  newly NotReady node. A newly appeared problem must never be absorbed into a\n  stale audit; it always gets a fresh investigation.\n- **Proceed** if there is no prior audit to compare against.\n\nNever link a periodic audit to an investigation that is still `PENDING_TRIAGE`,\n`PENDING_START`, or `IN_PROGRESS` — linking re-activates it and can keep it\nrunning indefinitely. Only skip (concurrent) or link to a completed one.\n\n## Notes\n\n- These are correlation preferences for the triage stage; the full diagnosis\n  (timeline, verdict, recommended actions) is produced later by the\n  `hyperpod-incident-rca` skill on investigations that proceed.\n- Routine `Info`-level activity and scale-in-progress churn (events mentioning\n  \"lost orchestration-ready status\") are already filtered upstream and should\n  not, on their own, drive a new investigation.\n",{"data":49,"body":54},{"name":4,"description":6,"metadata":50},{"version":51,"agent_types":52},"0.7.0",[53],"INCIDENT_TRIAGE",{"type":55,"children":56},"root",[57,65,86,91,139,146,172,274,279,285,297,383,412,418],{"type":58,"tag":59,"props":60,"children":61},"element","h1",{"id":4},[62],{"type":63,"value":64},"text","HyperPod incident triage",{"type":58,"tag":66,"props":67,"children":68},"p",{},[69,71,77,79,84],{"type":63,"value":70},"Use these rules when deciding whether a HyperPod incident should be ",{"type":58,"tag":72,"props":73,"children":74},"strong",{},[75],{"type":63,"value":76},"linked",{"type":63,"value":78},"\nto an existing investigation, ",{"type":58,"tag":72,"props":80,"children":81},{},[82],{"type":63,"value":83},"skipped",{"type":63,"value":85},", or investigated on its own. They\nrefine the default correlation for HyperPod's fault model. When a rule below\ndoes not clearly apply, fall back to your normal triage judgment.",{"type":58,"tag":66,"props":87,"children":88},{},[89],{"type":63,"value":90},"Every HyperPod incident carries a description built by the webhook bridge. Two\nkinds arrive:",{"type":58,"tag":92,"props":93,"children":94},"ul",{},[95,123],{"type":58,"tag":96,"props":97,"children":98},"li",{},[99,104,106,113,115,121],{"type":58,"tag":72,"props":100,"children":101},{},[102],{"type":63,"value":103},"Cluster events",{"type":63,"value":105}," (a real HyperPod fault) — the description contains the\ninstance group, a ",{"type":58,"tag":107,"props":108,"children":110},"code",{"className":109},[],[111],{"type":63,"value":112},"Description:",{"type":63,"value":114},", and often a ",{"type":58,"tag":107,"props":116,"children":118},{"className":117},[],[119],{"type":63,"value":120},"FailureMessage:",{"type":63,"value":122},".",{"type":58,"tag":96,"props":124,"children":125},{},[126,131,133],{"type":58,"tag":72,"props":127,"children":128},{},[129],{"type":63,"value":130},"Periodic audits",{"type":63,"value":132}," (a scheduled health sweep) — the description begins with\neither \"Periodic audit for HyperPod cluster '",{"type":58,"tag":134,"props":135,"children":136},"name",{},[137],{"type":63,"value":138},"' detected N issue(s)\"\n(issues found) or \"Daily heartbeat audit for HyperPod cluster\" (the once-daily\nall-clear liveness signal). Both ask the RCA skill to run in audit mode.",{"type":58,"tag":140,"props":141,"children":143},"h2",{"id":142},"linking-rules-for-hyperpod-fault-events",[144],{"type":63,"value":145},"Linking rules for HyperPod fault events",{"type":58,"tag":66,"props":147,"children":148},{},[149,151,156,158,163,165,170],{"type":63,"value":150},"Link a new HyperPod fault only to an existing HyperPod investigation that\ndescribes ",{"type":58,"tag":72,"props":152,"children":153},{},[154],{"type":63,"value":155},"the same fault on the same component",{"type":63,"value":157},". Treat the instance group\nplus the fault text (",{"type":58,"tag":107,"props":159,"children":161},{"className":160},[],[162],{"type":63,"value":112},{"type":63,"value":164}," together with any ",{"type":58,"tag":107,"props":166,"children":168},{"className":167},[],[169],{"type":63,"value":120},{"type":63,"value":171},") as the\nidentity of the fault.",{"type":58,"tag":92,"props":173,"children":174},{},[175,185,217,241],{"type":58,"tag":96,"props":176,"children":177},{},[178,183],{"type":58,"tag":72,"props":179,"children":180},{},[181],{"type":63,"value":182},"Do link",{"type":63,"value":184}," when both the instance group and the fault text match an open\ninvestigation — it is the same problem arriving again.",{"type":58,"tag":96,"props":186,"children":187},{},[188,193,195,201,203,208,210,215],{"type":58,"tag":72,"props":189,"children":190},{},[191],{"type":63,"value":192},"Do NOT link",{"type":63,"value":194}," when the instance group is the same but the fault text\ndiffers. Different fault types on the same instance group have different root\ncauses and need separate investigations. For example, a lifecycle-script\nfailure on ",{"type":58,"tag":107,"props":196,"children":198},{"className":197},[],[199],{"type":63,"value":200},"worker2",{"type":63,"value":202}," is ",{"type":58,"tag":72,"props":204,"children":205},{},[206],{"type":63,"value":207},"not",{"type":63,"value":209}," the same incident as a GPU\u002FXid fault on\n",{"type":58,"tag":107,"props":211,"children":213},{"className":212},[],[214],{"type":63,"value":200},{"type":63,"value":216}," — keep them separate even though they share the instance group.",{"type":58,"tag":96,"props":218,"children":219},{},[220,224,226,231,233,239],{"type":58,"tag":72,"props":221,"children":222},{},[223],{"type":63,"value":192},{"type":63,"value":225}," the same fault type across different instance groups (e.g.\nan Xid fault on ",{"type":58,"tag":107,"props":227,"children":229},{"className":228},[],[230],{"type":63,"value":200},{"type":63,"value":232}," and an Xid fault on ",{"type":58,"tag":107,"props":234,"children":236},{"className":235},[],[237],{"type":63,"value":238},"worker3",{"type":63,"value":240}," are two distinct\nhardware faults).",{"type":58,"tag":96,"props":242,"children":243},{},[244,248,250,256,258,264,266,272],{"type":58,"tag":72,"props":245,"children":246},{},[247],{"type":63,"value":192},{"type":63,"value":249}," capacity errors for different instance types (e.g.\n",{"type":58,"tag":107,"props":251,"children":253},{"className":252},[],[254],{"type":63,"value":255},"ml.g5.8xlarge",{"type":63,"value":257}," vs ",{"type":58,"tag":107,"props":259,"children":261},{"className":260},[],[262],{"type":63,"value":263},"ml.p5.48xlarge",{"type":63,"value":265},") — the ",{"type":58,"tag":107,"props":267,"children":269},{"className":268},[],[270],{"type":63,"value":271},"FailureMessage",{"type":63,"value":273}," differs, so they\nare distinct problems.",{"type":58,"tag":66,"props":275,"children":276},{},[277],{"type":63,"value":278},"When in doubt, prefer a separate investigation over an incorrect link:\nover-linking hides distinct problems, which is the failure this skill exists to\nprevent.",{"type":58,"tag":140,"props":280,"children":282},{"id":281},"rules-for-periodic-audits",[283],{"type":63,"value":284},"Rules for periodic audits",{"type":58,"tag":66,"props":286,"children":287},{},[288,290,295],{"type":63,"value":289},"A periodic audit is a scheduled sweep, not a specific fault. Decide based on\nwhether the cluster's condition has ",{"type":58,"tag":72,"props":291,"children":292},{},[293],{"type":63,"value":294},"changed",{"type":63,"value":296}," since the last audit:",{"type":58,"tag":92,"props":298,"children":299},{},[300,317,333,357,374],{"type":58,"tag":96,"props":301,"children":302},{},[303,308,310,315],{"type":58,"tag":72,"props":304,"children":305},{},[306],{"type":63,"value":307},"Skip",{"type":63,"value":309}," the ",{"type":58,"tag":72,"props":311,"children":312},{},[313],{"type":63,"value":314},"daily heartbeat \u002F all-clear audit",{"type":63,"value":316}," (description begins \"Daily\nheartbeat audit for HyperPod cluster\", or otherwise reports no open issues) —\nit carries no fault to investigate; the task's presence in the console is\nitself the liveness signal.",{"type":58,"tag":96,"props":318,"children":319},{},[320,324,326,331],{"type":58,"tag":72,"props":321,"children":322},{},[323],{"type":63,"value":307},{"type":63,"value":325}," the audit if another periodic-audit investigation for the same\ncluster is ",{"type":58,"tag":72,"props":327,"children":328},{},[329],{"type":63,"value":330},"currently in progress",{"type":63,"value":332}," — let it finish rather than starting or\nre-activating a parallel one.",{"type":58,"tag":96,"props":334,"children":335},{},[336,341,343,348,350,355],{"type":58,"tag":72,"props":337,"children":338},{},[339],{"type":63,"value":340},"Link",{"type":63,"value":342}," the audit to the most recent ",{"type":58,"tag":72,"props":344,"children":345},{},[346],{"type":63,"value":347},"completed",{"type":63,"value":349}," periodic-audit\ninvestigation only if the cluster's open problems are ",{"type":58,"tag":72,"props":351,"children":352},{},[353],{"type":63,"value":354},"unchanged",{"type":63,"value":356}," — the\nsame fault chains, the same CrashLoopBackOff pods, and the same NotReady\nnodes as that prior audit found. There is nothing new to report.",{"type":58,"tag":96,"props":358,"children":359},{},[360,365,367,372],{"type":58,"tag":72,"props":361,"children":362},{},[363],{"type":63,"value":364},"Proceed",{"type":63,"value":366}," (investigate) if the cluster shows ",{"type":58,"tag":72,"props":368,"children":369},{},[370],{"type":63,"value":371},"any new or changed problem",{"type":63,"value":373},"\nsince the last audit — a new fault event, a newly CrashLooping pod, or a\nnewly NotReady node. A newly appeared problem must never be absorbed into a\nstale audit; it always gets a fresh investigation.",{"type":58,"tag":96,"props":375,"children":376},{},[377,381],{"type":58,"tag":72,"props":378,"children":379},{},[380],{"type":63,"value":364},{"type":63,"value":382}," if there is no prior audit to compare against.",{"type":58,"tag":66,"props":384,"children":385},{},[386,388,394,396,402,404,410],{"type":63,"value":387},"Never link a periodic audit to an investigation that is still ",{"type":58,"tag":107,"props":389,"children":391},{"className":390},[],[392],{"type":63,"value":393},"PENDING_TRIAGE",{"type":63,"value":395},",\n",{"type":58,"tag":107,"props":397,"children":399},{"className":398},[],[400],{"type":63,"value":401},"PENDING_START",{"type":63,"value":403},", or ",{"type":58,"tag":107,"props":405,"children":407},{"className":406},[],[408],{"type":63,"value":409},"IN_PROGRESS",{"type":63,"value":411}," — linking re-activates it and can keep it\nrunning indefinitely. Only skip (concurrent) or link to a completed one.",{"type":58,"tag":140,"props":413,"children":415},{"id":414},"notes",[416],{"type":63,"value":417},"Notes",{"type":58,"tag":92,"props":419,"children":420},{},[421,434],{"type":58,"tag":96,"props":422,"children":423},{},[424,426,432],{"type":63,"value":425},"These are correlation preferences for the triage stage; the full diagnosis\n(timeline, verdict, recommended actions) is produced later by the\n",{"type":58,"tag":107,"props":427,"children":429},{"className":428},[],[430],{"type":63,"value":431},"hyperpod-incident-rca",{"type":63,"value":433}," skill on investigations that proceed.",{"type":58,"tag":96,"props":435,"children":436},{},[437,439,445],{"type":63,"value":438},"Routine ",{"type":58,"tag":107,"props":440,"children":442},{"className":441},[],[443],{"type":63,"value":444},"Info",{"type":63,"value":446},"-level activity and scale-in-progress churn (events mentioning\n\"lost orchestration-ready status\") are already filtered upstream and should\nnot, on their own, drive a new investigation.",{"items":448,"total":625},[449,468,489,499,512,525,535,545,563,574,594,610],{"slug":450,"name":450,"fn":451,"description":452,"org":453,"tags":454,"stars":465,"repoUrl":466,"updatedAt":467},"agentcore-investigation","investigate Bedrock AgentCore runtime sessions","Investigate Bedrock AgentCore runtime sessions via CloudWatch Logs Insights — resolve session\u002Ftrace IDs, query OTEL spans, filter noise, build timelines. Use when debugging AgentCore agent sessions, tracing tool calls, or analyzing latency.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[455,456,459,462],{"name":24,"slug":25,"type":16},{"name":457,"slug":458,"type":16},"Debugging","debugging",{"name":460,"slug":461,"type":16},"Logs","logs",{"name":463,"slug":464,"type":16},"Observability","observability",9645,"https:\u002F\u002Fgithub.com\u002Fawslabs\u002Fmcp","2026-07-12T08:37:22.601527",{"slug":469,"name":470,"fn":471,"description":472,"org":473,"tags":474,"stars":465,"repoUrl":466,"updatedAt":488},"amazon-aurora-dsql","amazon aurora dsql","build applications with Aurora DSQL","Deprecated compatibility redirect for Aurora DSQL guidance. Use when a request concerns DSQL, Aurora DSQL, distributed SQL, DSQL schemas, migrations, queries, authentication, performance, or application development.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[475,478,479,482,485],{"name":476,"slug":477,"type":16},"Aurora","aurora",{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},"Database","database",{"name":483,"slug":484,"type":16},"Serverless","serverless",{"name":486,"slug":487,"type":16},"SQL","sql","2026-09-02T07:20:51.53702",{"slug":490,"name":491,"fn":471,"description":472,"org":492,"tags":493,"stars":465,"repoUrl":466,"updatedAt":498},"aurora-dsql","aurora dsql",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[494,495,496,497],{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},{"name":483,"slug":484,"type":16},{"name":486,"slug":487,"type":16},"2026-09-02T07:20:46.533217",{"slug":500,"name":501,"fn":471,"description":472,"org":502,"tags":503,"stars":465,"repoUrl":466,"updatedAt":511},"aws-dsql","aws dsql",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[504,505,506,509,510],{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},{"name":507,"slug":508,"type":16},"Migration","migration",{"name":483,"slug":484,"type":16},{"name":486,"slug":487,"type":16},"2026-09-02T07:20:49.531712",{"slug":513,"name":514,"fn":471,"description":472,"org":515,"tags":516,"stars":465,"repoUrl":466,"updatedAt":524},"distributed-postgres","distributed postgres",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[517,518,519,522,523],{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},{"name":520,"slug":521,"type":16},"PostgreSQL","postgresql",{"name":483,"slug":484,"type":16},{"name":486,"slug":487,"type":16},"2026-09-02T07:20:47.592534",{"slug":526,"name":527,"fn":471,"description":472,"org":528,"tags":529,"stars":465,"repoUrl":466,"updatedAt":534},"distributed-sql","distributed sql",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[530,531,532,533],{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},{"name":483,"slug":484,"type":16},{"name":486,"slug":487,"type":16},"2026-09-02T07:20:50.520015",{"slug":536,"name":536,"fn":471,"description":472,"org":537,"tags":538,"stars":465,"repoUrl":466,"updatedAt":544},"dsql",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[539,540,541,542,543],{"name":24,"slug":25,"type":16},{"name":480,"slug":481,"type":16},{"name":507,"slug":508,"type":16},{"name":483,"slug":484,"type":16},{"name":486,"slug":487,"type":16},"2026-09-02T07:20:48.570617",{"slug":546,"name":546,"fn":547,"description":548,"org":549,"tags":550,"stars":560,"repoUrl":561,"updatedAt":562},"aidlc","orchestrate AI-driven development lifecycle workflows","AI-DLC workflow orchestrator. Start, resume, or manage an AI-driven development lifecycle. Scopes are defined one file per scope under `.kiro\u002Fscopes\u002F`; run `bun .kiro\u002Ftools\u002Faidlc-utility.ts help` for the authoritative list and descriptions. Utilities: --status, --doctor, --stage, --phase, --scope, --depth, --test-strategy, --review, --version, --help, plus the intent and space verbs. Or describe what you want to build and the scope will be auto-detected.\n",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[551,554,557],{"name":552,"slug":553,"type":16},"Agents","agents",{"name":555,"slug":556,"type":16},"Automation","automation",{"name":558,"slug":559,"type":16},"Workflow","workflow",4261,"https:\u002F\u002Fgithub.com\u002Fawslabs\u002Faidlc-workflows","2026-09-02T07:47:34.75815",{"slug":564,"name":564,"fn":565,"description":566,"org":567,"tags":568,"stars":560,"repoUrl":561,"updatedAt":573},"aidlc-jump","navigate AI-DLC workflow stages and phases","Jump the active AI-DLC workflow to a stage or phase. A Cursor-native shortcut for `\u002Faidlc --stage \u003Ctarget>` or `\u002Faidlc --phase \u003Ctarget>`.\n",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[569,570],{"name":555,"slug":556,"type":16},{"name":571,"slug":572,"type":16},"Navigation","navigation","2026-09-02T07:48:09.505949",{"slug":575,"name":575,"fn":576,"description":577,"org":578,"tags":579,"stars":560,"repoUrl":561,"updatedAt":593},"aidlc-knowledge","index documents for AI-DLC agent citation","Index the team's own documents — PDFs, Word files, Markdown, plain text — into a per-space catalog the AI-DLC agents can cite. Wraps `aidlc-knowledge.ts`: onboard, sync, list, show, associate, dissociate, rebind, summarize. Every catalog row is written by the tool under a workspace lock; this skill never edits the catalog by hand and never advances workflow state.\n",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[580,581,584,587,590],{"name":552,"slug":553,"type":16},{"name":582,"slug":583,"type":16},"Documents","documents",{"name":585,"slug":586,"type":16},"Knowledge Management","knowledge-management",{"name":588,"slug":589,"type":16},"Markdown","markdown",{"name":591,"slug":592,"type":16},"PDF","pdf","2026-09-02T07:48:10.614775",{"slug":595,"name":595,"fn":596,"description":597,"org":598,"tags":599,"stars":560,"repoUrl":561,"updatedAt":609},"aidlc-outcomes-pack","generate AI-DLC workflow handover documentation","Generate a comprehensive handover document at workflow close so the team can own, operate, and continue the system without re-running the workflow. Stage\u002Fphase\u002Flearning counts come from `aidlc-runtime.ts summary`; prose comes from the artefacts. Writes OUTCOMES.md but never mutates workflow state or emits audit events.\n",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[600,603,606],{"name":601,"slug":602,"type":16},"Documentation","documentation",{"name":604,"slug":605,"type":16},"Process Documentation","process-documentation",{"name":607,"slug":608,"type":16},"Reporting","reporting","2026-09-02T07:47:34.212738",{"slug":611,"name":611,"fn":612,"description":613,"org":614,"tags":615,"stars":560,"repoUrl":561,"updatedAt":624},"aidlc-replay","generate AI-DLC session narrative reports","Print a structured session narrative for stakeholders who weren't in the room. Numbers (stage counts, phase rollup, duration) come from `aidlc-runtime.ts summary`; prose comes from the audit trail and artefacts. Renders to the terminal only — writes no file, never mutates workflow state, never emits audit events.\n",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[616,619,620,623],{"name":617,"slug":618,"type":16},"Audit","audit",{"name":601,"slug":602,"type":16},{"name":621,"slug":622,"type":16},"Engineering","engineering",{"name":607,"slug":608,"type":16},"2026-09-02T07:47:45.055994",142,{"items":627,"total":661},[628,644,654],{"slug":629,"name":629,"fn":630,"description":631,"org":632,"tags":633,"stars":26,"repoUrl":27,"updatedAt":643},"hyperpod-devops-agent-solution","monitor SageMaker HyperPod infrastructure","How this Agent Space monitors SageMaker HyperPod — the design and intent of the HyperPod x DevOps Agent solution (event-driven webhook bridge, Lambda-gated periodic audit, triage\u002FRCA skills, email notifications). Read this to understand WHY a HyperPod investigation was created and what the monitoring pipeline does. For the concrete resource\u002Ftopology map (ARNs, IDs, log groups), see the understanding-agent-space skill instead.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[634,635,636,639,642],{"name":552,"slug":553,"type":16},{"name":24,"slug":25,"type":16},{"name":637,"slug":638,"type":16},"DevOps","devops",{"name":640,"slug":641,"type":16},"Monitoring","monitoring",{"name":463,"slug":464,"type":16},"2026-09-02T07:47:35.452244",{"slug":431,"name":431,"fn":645,"description":646,"org":647,"tags":648,"stars":26,"repoUrl":27,"updatedAt":653},"perform root-cause analysis for HyperPod incidents","Root-cause analysis for a SageMaker HyperPod incident, after triage has decided to PROCEED. Runs at the INCIDENT_RCA stage. Reads describe-cluster, list-cluster-nodes, list-cluster-events, and HMA CloudWatch streams; reconstructs a timeline; classifies as Suppress \u002F Monitor \u002F Escalate \u002F Resolved against time budgets and recurrence statistics from the HyperPod mental model. Produces a human-readable verdict report with recommended operator actions. The complementary INCIDENT_TRIAGE skill `hyperpod-incident-triage` decides LINKED \u002F SKIPPED \u002F PROCEED before this skill runs.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[649,650,651,652],{"name":24,"slug":25,"type":16},{"name":457,"slug":458,"type":16},{"name":18,"slug":19,"type":16},{"name":14,"slug":15,"type":16},"2026-09-02T07:47:38.297082",{"slug":4,"name":4,"fn":5,"description":6,"org":655,"tags":656,"stars":26,"repoUrl":27,"updatedAt":28},{"slug":8,"name":9,"logoUrl":10,"githubOrg":11},[657,658,659,660],{"name":24,"slug":25,"type":16},{"name":18,"slug":19,"type":16},{"name":14,"slug":15,"type":16},{"name":21,"slug":22,"type":16},3]