[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"skill-sourcegraph-run":3,"mdc-hdt2ao-key":32,"related-repo-sourcegraph-run":668,"related-org-sourcegraph-run":757},{"slug":4,"name":4,"fn":5,"description":6,"org":7,"tags":11,"stars":22,"repoUrl":23,"updatedAt":24,"license":25,"forks":26,"topics":27,"repo":28,"sourceUrl":30,"mdContent":31},"run","manage CodeScaleBench benchmark runs","Launch and manage CodeScaleBench benchmark runs with paired-run guardrails, quick reruns, and execution orchestration.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},"sourcegraph","Sourcegraph","https:\u002F\u002Fpexgzepcugksgbtrxkhf.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Forg-logos\u002Fsourcegraph.png",[12,16,19],{"name":13,"slug":14,"type":15},"Benchmarking","benchmarking","tag",{"name":17,"slug":18,"type":15},"Automation","automation",{"name":20,"slug":21,"type":15},"Testing","testing",31,"https:\u002F\u002Fgithub.com\u002Fsourcegraph\u002FCodeScaleBench","2026-07-16T06:04:40.848817",null,4,[],{"repoUrl":23,"stars":22,"forks":26,"topics":29,"description":25},[],"https:\u002F\u002Fgithub.com\u002Fsourcegraph\u002FCodeScaleBench\u002Ftree\u002FHEAD\u002Fskills\u002Frun","---\nname: run\ndescription: Launch and manage CodeScaleBench benchmark runs with paired-run guardrails, quick reruns, and execution orchestration.\n---\n\n# Skill: Run Benchmarks\n\n## Scope\n\nUse this skill when the user asks to:\n- Run benchmark suites, rerun failures, or launch gap-fill batches\n- Manage multi-account parallel execution\n- Execute paired baseline+MCP runs with curation guardrails\n- Perform quick reruns of specific tasks or suites\n\n## Approval Gate (Required Before Running)\n\nBefore executing any benchmark run, confirm with the user:\n\n1. **Model** — which model? (e.g., `anthropic\u002Fclaude-haiku-4-5-20251001` for test runs)\n2. **Suite \u002F selection file** — which benchmark suite or `--selection-file`?\n3. **Config** — paired (default), `--baseline-only`, or `--full-only`? Which `--full-config`?\n4. **Parallel slots** — how many? (default: auto-detect; use 8+ for multi-account)\n5. **Category** — `staging` (default) or `official`?\n\n**Do NOT launch a run until the user has confirmed these five parameters.**\n\n## Canonical Commands\n\n- Per-suite default: `.\u002Fconfigs\u002Fharnesses\u002F\u003Csuite>_2config.sh`\n- Unified selected-task runner: `.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh`\n- Config registry: `configs\u002Feval_matrix.json`\n- Quick rerun: use `--rerun` flag on base command\n\n## Run Policy (Mandatory)\n\n- Default execution is **paired by task**: `baseline` + `sourcegraph_full`\n- Single-lane runs are **gap-fill only**:\n  - `--baseline-only` requires valid existing `sourcegraph_full` counterpart runs\n  - `--full-only` requires valid existing `baseline` counterpart runs\n- Emergency bypass only: `ALLOW_UNPAIRED_SINGLE_CONFIG=true`\n- Account readiness: Always run `python3 scripts\u002Finfra\u002Faccount_health.py status` before launching\n\n## Standard Launch Patterns\n\n```bash\n# Paired per-suite run\n.\u002Fconfigs\u002Fharnesses\u002Fpytorch_2config.sh --parallel 4\n\n# Paired selected-task run\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch\n\n# Gap-fill baseline only (guarded)\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch --baseline-only\n\n# Quick rerun of failed tasks\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch --rerun failed\n```\n\n## Infrastructure\n\n- **Default environment**: Daytona (see `docs\u002FDAYTONA.md`)\n- **Parallelism**: Auto-detected from account count and rate limits; override with `--parallel N`\n- **Orchestration**: `scripts\u002Frunning\u002Fcontrol_plane.py` manages multi-account scheduling\n- **Monitoring**: `scripts\u002Frunning\u002Fmonitor_and_queue.sh` watches active runs\n\n## Related Skills\n\n- `\u002Fstatus` — monitor active runs, check completion\n- `\u002Faudit` — post-run validation and integrity checks\n- `\u002Fevaluate` — extract and score results\n",{"data":33,"body":34},{"name":4,"description":6},{"type":35,"children":36},"root",[37,46,53,59,84,90,95,205,213,219,268,274,374,380,543,549,620,626,662],{"type":38,"tag":39,"props":40,"children":42},"element","h1",{"id":41},"skill-run-benchmarks",[43],{"type":44,"value":45},"text","Skill: Run Benchmarks",{"type":38,"tag":47,"props":48,"children":50},"h2",{"id":49},"scope",[51],{"type":44,"value":52},"Scope",{"type":38,"tag":54,"props":55,"children":56},"p",{},[57],{"type":44,"value":58},"Use this skill when the user asks to:",{"type":38,"tag":60,"props":61,"children":62},"ul",{},[63,69,74,79],{"type":38,"tag":64,"props":65,"children":66},"li",{},[67],{"type":44,"value":68},"Run benchmark suites, rerun failures, or launch gap-fill batches",{"type":38,"tag":64,"props":70,"children":71},{},[72],{"type":44,"value":73},"Manage multi-account parallel execution",{"type":38,"tag":64,"props":75,"children":76},{},[77],{"type":44,"value":78},"Execute paired baseline+MCP runs with curation guardrails",{"type":38,"tag":64,"props":80,"children":81},{},[82],{"type":44,"value":83},"Perform quick reruns of specific tasks or suites",{"type":38,"tag":47,"props":85,"children":87},{"id":86},"approval-gate-required-before-running",[88],{"type":44,"value":89},"Approval Gate (Required Before Running)",{"type":38,"tag":54,"props":91,"children":92},{},[93],{"type":44,"value":94},"Before executing any benchmark run, confirm with the user:",{"type":38,"tag":96,"props":97,"children":98},"ol",{},[99,119,137,170,180],{"type":38,"tag":64,"props":100,"children":101},{},[102,108,110,117],{"type":38,"tag":103,"props":104,"children":105},"strong",{},[106],{"type":44,"value":107},"Model",{"type":44,"value":109}," — which model? (e.g., ",{"type":38,"tag":111,"props":112,"children":114},"code",{"className":113},[],[115],{"type":44,"value":116},"anthropic\u002Fclaude-haiku-4-5-20251001",{"type":44,"value":118}," for test runs)",{"type":38,"tag":64,"props":120,"children":121},{},[122,127,129,135],{"type":38,"tag":103,"props":123,"children":124},{},[125],{"type":44,"value":126},"Suite \u002F selection file",{"type":44,"value":128}," — which benchmark suite or ",{"type":38,"tag":111,"props":130,"children":132},{"className":131},[],[133],{"type":44,"value":134},"--selection-file",{"type":44,"value":136},"?",{"type":38,"tag":64,"props":138,"children":139},{},[140,145,147,153,155,161,163,169],{"type":38,"tag":103,"props":141,"children":142},{},[143],{"type":44,"value":144},"Config",{"type":44,"value":146}," — paired (default), ",{"type":38,"tag":111,"props":148,"children":150},{"className":149},[],[151],{"type":44,"value":152},"--baseline-only",{"type":44,"value":154},", or ",{"type":38,"tag":111,"props":156,"children":158},{"className":157},[],[159],{"type":44,"value":160},"--full-only",{"type":44,"value":162},"? Which ",{"type":38,"tag":111,"props":164,"children":166},{"className":165},[],[167],{"type":44,"value":168},"--full-config",{"type":44,"value":136},{"type":38,"tag":64,"props":171,"children":172},{},[173,178],{"type":38,"tag":103,"props":174,"children":175},{},[176],{"type":44,"value":177},"Parallel slots",{"type":44,"value":179}," — how many? (default: auto-detect; use 8+ for multi-account)",{"type":38,"tag":64,"props":181,"children":182},{},[183,188,190,196,198,204],{"type":38,"tag":103,"props":184,"children":185},{},[186],{"type":44,"value":187},"Category",{"type":44,"value":189}," — ",{"type":38,"tag":111,"props":191,"children":193},{"className":192},[],[194],{"type":44,"value":195},"staging",{"type":44,"value":197}," (default) or ",{"type":38,"tag":111,"props":199,"children":201},{"className":200},[],[202],{"type":44,"value":203},"official",{"type":44,"value":136},{"type":38,"tag":54,"props":206,"children":207},{},[208],{"type":38,"tag":103,"props":209,"children":210},{},[211],{"type":44,"value":212},"Do NOT launch a run until the user has confirmed these five parameters.",{"type":38,"tag":47,"props":214,"children":216},{"id":215},"canonical-commands",[217],{"type":44,"value":218},"Canonical Commands",{"type":38,"tag":60,"props":220,"children":221},{},[222,233,244,255],{"type":38,"tag":64,"props":223,"children":224},{},[225,227],{"type":44,"value":226},"Per-suite default: ",{"type":38,"tag":111,"props":228,"children":230},{"className":229},[],[231],{"type":44,"value":232},".\u002Fconfigs\u002Fharnesses\u002F\u003Csuite>_2config.sh",{"type":38,"tag":64,"props":234,"children":235},{},[236,238],{"type":44,"value":237},"Unified selected-task runner: ",{"type":38,"tag":111,"props":239,"children":241},{"className":240},[],[242],{"type":44,"value":243},".\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh",{"type":38,"tag":64,"props":245,"children":246},{},[247,249],{"type":44,"value":248},"Config registry: ",{"type":38,"tag":111,"props":250,"children":252},{"className":251},[],[253],{"type":44,"value":254},"configs\u002Feval_matrix.json",{"type":38,"tag":64,"props":256,"children":257},{},[258,260,266],{"type":44,"value":259},"Quick rerun: use ",{"type":38,"tag":111,"props":261,"children":263},{"className":262},[],[264],{"type":44,"value":265},"--rerun",{"type":44,"value":267}," flag on base command",{"type":38,"tag":47,"props":269,"children":271},{"id":270},"run-policy-mandatory",[272],{"type":44,"value":273},"Run Policy (Mandatory)",{"type":38,"tag":60,"props":275,"children":276},{},[277,303,350,361],{"type":38,"tag":64,"props":278,"children":279},{},[280,282,287,289,295,297],{"type":44,"value":281},"Default execution is ",{"type":38,"tag":103,"props":283,"children":284},{},[285],{"type":44,"value":286},"paired by task",{"type":44,"value":288},": ",{"type":38,"tag":111,"props":290,"children":292},{"className":291},[],[293],{"type":44,"value":294},"baseline",{"type":44,"value":296}," + ",{"type":38,"tag":111,"props":298,"children":300},{"className":299},[],[301],{"type":44,"value":302},"sourcegraph_full",{"type":38,"tag":64,"props":304,"children":305},{},[306,308,313,315],{"type":44,"value":307},"Single-lane runs are ",{"type":38,"tag":103,"props":309,"children":310},{},[311],{"type":44,"value":312},"gap-fill only",{"type":44,"value":314},":\n",{"type":38,"tag":60,"props":316,"children":317},{},[318,335],{"type":38,"tag":64,"props":319,"children":320},{},[321,326,328,333],{"type":38,"tag":111,"props":322,"children":324},{"className":323},[],[325],{"type":44,"value":152},{"type":44,"value":327}," requires valid existing ",{"type":38,"tag":111,"props":329,"children":331},{"className":330},[],[332],{"type":44,"value":302},{"type":44,"value":334}," counterpart runs",{"type":38,"tag":64,"props":336,"children":337},{},[338,343,344,349],{"type":38,"tag":111,"props":339,"children":341},{"className":340},[],[342],{"type":44,"value":160},{"type":44,"value":327},{"type":38,"tag":111,"props":345,"children":347},{"className":346},[],[348],{"type":44,"value":294},{"type":44,"value":334},{"type":38,"tag":64,"props":351,"children":352},{},[353,355],{"type":44,"value":354},"Emergency bypass only: ",{"type":38,"tag":111,"props":356,"children":358},{"className":357},[],[359],{"type":44,"value":360},"ALLOW_UNPAIRED_SINGLE_CONFIG=true",{"type":38,"tag":64,"props":362,"children":363},{},[364,366,372],{"type":44,"value":365},"Account readiness: Always run ",{"type":38,"tag":111,"props":367,"children":369},{"className":368},[],[370],{"type":44,"value":371},"python3 scripts\u002Finfra\u002Faccount_health.py status",{"type":44,"value":373}," before launching",{"type":38,"tag":47,"props":375,"children":377},{"id":376},"standard-launch-patterns",[378],{"type":44,"value":379},"Standard Launch Patterns",{"type":38,"tag":381,"props":382,"children":387},"pre",{"className":383,"code":384,"language":385,"meta":386,"style":386},"language-bash shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","# Paired per-suite run\n.\u002Fconfigs\u002Fharnesses\u002Fpytorch_2config.sh --parallel 4\n\n# Paired selected-task run\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch\n\n# Gap-fill baseline only (guarded)\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch --baseline-only\n\n# Quick rerun of failed tasks\n.\u002Fconfigs\u002Fharnesses\u002Frun_selected_tasks.sh --benchmark csb_sdlc_pytorch --rerun failed\n","bash","",[388],{"type":38,"tag":111,"props":389,"children":390},{"__ignoreMap":386},[391,403,425,435,443,461,469,478,500,508,517],{"type":38,"tag":392,"props":393,"children":396},"span",{"class":394,"line":395},"line",1,[397],{"type":38,"tag":392,"props":398,"children":400},{"style":399},"--shiki-light:#90A4AE;--shiki-light-font-style:italic;--shiki-default:#546E7A;--shiki-default-font-style:italic;--shiki-dark:#676E95;--shiki-dark-font-style:italic",[401],{"type":44,"value":402},"# Paired per-suite run\n",{"type":38,"tag":392,"props":404,"children":406},{"class":394,"line":405},2,[407,413,419],{"type":38,"tag":392,"props":408,"children":410},{"style":409},"--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B",[411],{"type":44,"value":412},".\u002Fconfigs\u002Fharnesses\u002Fpytorch_2config.sh",{"type":38,"tag":392,"props":414,"children":416},{"style":415},"--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D",[417],{"type":44,"value":418}," --parallel",{"type":38,"tag":392,"props":420,"children":422},{"style":421},"--shiki-light:#F76D47;--shiki-default:#F78C6C;--shiki-dark:#F78C6C",[423],{"type":44,"value":424}," 4\n",{"type":38,"tag":392,"props":426,"children":428},{"class":394,"line":427},3,[429],{"type":38,"tag":392,"props":430,"children":432},{"emptyLinePlaceholder":431},true,[433],{"type":44,"value":434},"\n",{"type":38,"tag":392,"props":436,"children":437},{"class":394,"line":26},[438],{"type":38,"tag":392,"props":439,"children":440},{"style":399},[441],{"type":44,"value":442},"# Paired selected-task run\n",{"type":38,"tag":392,"props":444,"children":446},{"class":394,"line":445},5,[447,451,456],{"type":38,"tag":392,"props":448,"children":449},{"style":409},[450],{"type":44,"value":243},{"type":38,"tag":392,"props":452,"children":453},{"style":415},[454],{"type":44,"value":455}," --benchmark",{"type":38,"tag":392,"props":457,"children":458},{"style":415},[459],{"type":44,"value":460}," csb_sdlc_pytorch\n",{"type":38,"tag":392,"props":462,"children":464},{"class":394,"line":463},6,[465],{"type":38,"tag":392,"props":466,"children":467},{"emptyLinePlaceholder":431},[468],{"type":44,"value":434},{"type":38,"tag":392,"props":470,"children":472},{"class":394,"line":471},7,[473],{"type":38,"tag":392,"props":474,"children":475},{"style":399},[476],{"type":44,"value":477},"# Gap-fill baseline only (guarded)\n",{"type":38,"tag":392,"props":479,"children":481},{"class":394,"line":480},8,[482,486,490,495],{"type":38,"tag":392,"props":483,"children":484},{"style":409},[485],{"type":44,"value":243},{"type":38,"tag":392,"props":487,"children":488},{"style":415},[489],{"type":44,"value":455},{"type":38,"tag":392,"props":491,"children":492},{"style":415},[493],{"type":44,"value":494}," csb_sdlc_pytorch",{"type":38,"tag":392,"props":496,"children":497},{"style":415},[498],{"type":44,"value":499}," --baseline-only\n",{"type":38,"tag":392,"props":501,"children":503},{"class":394,"line":502},9,[504],{"type":38,"tag":392,"props":505,"children":506},{"emptyLinePlaceholder":431},[507],{"type":44,"value":434},{"type":38,"tag":392,"props":509,"children":511},{"class":394,"line":510},10,[512],{"type":38,"tag":392,"props":513,"children":514},{"style":399},[515],{"type":44,"value":516},"# Quick rerun of failed tasks\n",{"type":38,"tag":392,"props":518,"children":520},{"class":394,"line":519},11,[521,525,529,533,538],{"type":38,"tag":392,"props":522,"children":523},{"style":409},[524],{"type":44,"value":243},{"type":38,"tag":392,"props":526,"children":527},{"style":415},[528],{"type":44,"value":455},{"type":38,"tag":392,"props":530,"children":531},{"style":415},[532],{"type":44,"value":494},{"type":38,"tag":392,"props":534,"children":535},{"style":415},[536],{"type":44,"value":537}," --rerun",{"type":38,"tag":392,"props":539,"children":540},{"style":415},[541],{"type":44,"value":542}," failed\n",{"type":38,"tag":47,"props":544,"children":546},{"id":545},"infrastructure",[547],{"type":44,"value":548},"Infrastructure",{"type":38,"tag":60,"props":550,"children":551},{},[552,570,586,603],{"type":38,"tag":64,"props":553,"children":554},{},[555,560,562,568],{"type":38,"tag":103,"props":556,"children":557},{},[558],{"type":44,"value":559},"Default environment",{"type":44,"value":561},": Daytona (see ",{"type":38,"tag":111,"props":563,"children":565},{"className":564},[],[566],{"type":44,"value":567},"docs\u002FDAYTONA.md",{"type":44,"value":569},")",{"type":38,"tag":64,"props":571,"children":572},{},[573,578,580],{"type":38,"tag":103,"props":574,"children":575},{},[576],{"type":44,"value":577},"Parallelism",{"type":44,"value":579},": Auto-detected from account count and rate limits; override with ",{"type":38,"tag":111,"props":581,"children":583},{"className":582},[],[584],{"type":44,"value":585},"--parallel N",{"type":38,"tag":64,"props":587,"children":588},{},[589,594,595,601],{"type":38,"tag":103,"props":590,"children":591},{},[592],{"type":44,"value":593},"Orchestration",{"type":44,"value":288},{"type":38,"tag":111,"props":596,"children":598},{"className":597},[],[599],{"type":44,"value":600},"scripts\u002Frunning\u002Fcontrol_plane.py",{"type":44,"value":602}," manages multi-account scheduling",{"type":38,"tag":64,"props":604,"children":605},{},[606,611,612,618],{"type":38,"tag":103,"props":607,"children":608},{},[609],{"type":44,"value":610},"Monitoring",{"type":44,"value":288},{"type":38,"tag":111,"props":613,"children":615},{"className":614},[],[616],{"type":44,"value":617},"scripts\u002Frunning\u002Fmonitor_and_queue.sh",{"type":44,"value":619}," watches active runs",{"type":38,"tag":47,"props":621,"children":623},{"id":622},"related-skills",[624],{"type":44,"value":625},"Related Skills",{"type":38,"tag":60,"props":627,"children":628},{},[629,640,651],{"type":38,"tag":64,"props":630,"children":631},{},[632,638],{"type":38,"tag":111,"props":633,"children":635},{"className":634},[],[636],{"type":44,"value":637},"\u002Fstatus",{"type":44,"value":639}," — monitor active runs, check completion",{"type":38,"tag":64,"props":641,"children":642},{},[643,649],{"type":38,"tag":111,"props":644,"children":646},{"className":645},[],[647],{"type":44,"value":648},"\u002Faudit",{"type":44,"value":650}," — post-run validation and integrity checks",{"type":38,"tag":64,"props":652,"children":653},{},[654,660],{"type":38,"tag":111,"props":655,"children":657},{"className":656},[],[658],{"type":44,"value":659},"\u002Fevaluate",{"type":44,"value":661}," — extract and score results",{"type":38,"tag":663,"props":664,"children":665},"style",{},[666],{"type":44,"value":667},"html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"items":669,"total":502},[670,683,700,713,727,741,747],{"slug":671,"name":671,"fn":672,"description":673,"org":674,"tags":675,"stars":22,"repoUrl":23,"updatedAt":682},"audit","audit repository health and benchmark integrity","Run repo health checks, validate benchmark tasks, and audit run integrity.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[676,678,679],{"name":677,"slug":671,"type":15},"Audit",{"name":13,"slug":14,"type":15},{"name":680,"slug":681,"type":15},"QA","qa","2026-07-17T06:07:07.220218",{"slug":684,"name":684,"fn":685,"description":686,"org":687,"tags":688,"stars":22,"repoUrl":23,"updatedAt":699},"evaluate","score traces and evaluate benchmark results","Extract metrics, score traces, and evaluate benchmark task results.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[689,690,693,696],{"name":13,"slug":14,"type":15},{"name":691,"slug":692,"type":15},"Data Analysis","data-analysis",{"name":694,"slug":695,"type":15},"Engineering","engineering",{"name":697,"slug":698,"type":15},"Evals","evals","2026-07-16T06:04:41.910821",{"slug":701,"name":701,"fn":702,"description":703,"org":704,"tags":705,"stars":22,"repoUrl":23,"updatedAt":712},"infra","check infrastructure and system dependencies","Check infrastructure readiness, manage MCP tools, and audit system dependencies.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[706,707,708,709],{"name":677,"slug":671,"type":15},{"name":694,"slug":695,"type":15},{"name":548,"slug":545,"type":15},{"name":710,"slug":711,"type":15},"MCP","mcp","2026-07-16T06:04:41.188274",{"slug":714,"name":714,"fn":715,"description":716,"org":717,"tags":718,"stars":22,"repoUrl":23,"updatedAt":726},"next","plan benchmarking and coverage tasks","Plan upcoming work, analyze coverage gaps, and recommend next steps for benchmarking.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[719,720,723],{"name":13,"slug":14,"type":15},{"name":721,"slug":722,"type":15},"Planning","planning",{"name":724,"slug":725,"type":15},"Strategy","strategy","2026-07-17T06:06:57.69018",{"slug":728,"name":728,"fn":729,"description":730,"org":731,"tags":732,"stars":22,"repoUrl":23,"updatedAt":740},"report","generate CodeScaleBench evaluation reports","Generate evaluation reports, analyze run costs, and compare configurations.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[733,736,737],{"name":734,"slug":735,"type":15},"Analytics","analytics",{"name":13,"slug":14,"type":15},{"name":738,"slug":739,"type":15},"Reporting","reporting","2026-07-16T06:02:36.809556",{"slug":4,"name":4,"fn":5,"description":6,"org":742,"tags":743,"stars":22,"repoUrl":23,"updatedAt":24},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[744,745,746],{"name":17,"slug":18,"type":15},{"name":13,"slug":14,"type":15},{"name":20,"slug":21,"type":15},{"slug":748,"name":748,"fn":749,"description":750,"org":751,"tags":752,"stars":22,"repoUrl":23,"updatedAt":756},"scaffold","create and validate benchmark tasks","Create, mine, and validate new benchmark tasks and task suites.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[753,754,755],{"name":13,"slug":14,"type":15},{"name":694,"slug":695,"type":15},{"name":20,"slug":21,"type":15},"2026-07-16T06:04:41.525889",{"items":758,"total":502},[759,765,772,779,785,791,797,803,813],{"slug":671,"name":671,"fn":672,"description":673,"org":760,"tags":761,"stars":22,"repoUrl":23,"updatedAt":682},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[762,763,764],{"name":677,"slug":671,"type":15},{"name":13,"slug":14,"type":15},{"name":680,"slug":681,"type":15},{"slug":684,"name":684,"fn":685,"description":686,"org":766,"tags":767,"stars":22,"repoUrl":23,"updatedAt":699},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[768,769,770,771],{"name":13,"slug":14,"type":15},{"name":691,"slug":692,"type":15},{"name":694,"slug":695,"type":15},{"name":697,"slug":698,"type":15},{"slug":701,"name":701,"fn":702,"description":703,"org":773,"tags":774,"stars":22,"repoUrl":23,"updatedAt":712},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[775,776,777,778],{"name":677,"slug":671,"type":15},{"name":694,"slug":695,"type":15},{"name":548,"slug":545,"type":15},{"name":710,"slug":711,"type":15},{"slug":714,"name":714,"fn":715,"description":716,"org":780,"tags":781,"stars":22,"repoUrl":23,"updatedAt":726},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[782,783,784],{"name":13,"slug":14,"type":15},{"name":721,"slug":722,"type":15},{"name":724,"slug":725,"type":15},{"slug":728,"name":728,"fn":729,"description":730,"org":786,"tags":787,"stars":22,"repoUrl":23,"updatedAt":740},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[788,789,790],{"name":734,"slug":735,"type":15},{"name":13,"slug":14,"type":15},{"name":738,"slug":739,"type":15},{"slug":4,"name":4,"fn":5,"description":6,"org":792,"tags":793,"stars":22,"repoUrl":23,"updatedAt":24},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[794,795,796],{"name":17,"slug":18,"type":15},{"name":13,"slug":14,"type":15},{"name":20,"slug":21,"type":15},{"slug":748,"name":748,"fn":749,"description":750,"org":798,"tags":799,"stars":22,"repoUrl":23,"updatedAt":756},{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[800,801,802],{"name":13,"slug":14,"type":15},{"name":694,"slug":695,"type":15},{"name":20,"slug":21,"type":15},{"slug":804,"name":804,"fn":805,"description":806,"org":807,"tags":808,"stars":22,"repoUrl":23,"updatedAt":812},"status","monitor benchmark execution and task status","Monitor active runs, check task completion status, and watch benchmark execution progress.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[809,810],{"name":13,"slug":14,"type":15},{"name":610,"slug":811,"type":15},"monitoring","2026-07-16T06:02:37.137106",{"slug":814,"name":814,"fn":815,"description":816,"org":817,"tags":818,"stars":22,"repoUrl":23,"updatedAt":825},"triage","triage and analyze failed benchmark tasks","Investigate and triage failed benchmark tasks, analyze root causes, and plan reruns.",{"slug":8,"name":9,"logoUrl":10,"githubOrg":8},[819,820,823],{"name":13,"slug":14,"type":15},{"name":821,"slug":822,"type":15},"Debugging","debugging",{"name":824,"slug":814,"type":15},"Triage","2026-07-16T06:02:37.472337"]