
Skill
context-intelligence-evaluation-methodology
design evaluation metrics for tool signals
Description
Use when deciding how to measure a context-intelligence tool signal — metric design across quality/efficiency/efficacy axes, artifact-metric avoidance via precursor measurement, A/B and statistical-N discipline, and test-data fidelity.
SKILL.md
Context Intelligence Evaluation Methodology
Mode-only skill for how to measure. It complements — and never restates —
context-intelligence-eval-design (which owns scenario mechanics and the two-layer
structural/behavioral structure) and digital-twin-universe (which owns the DTU machinery).
Scope
In scope
- Metric design. Choose metrics across three axes: quality (did it detect what the user means?), efficiency (token/tool cost to detect), efficacy (does detection drive the right outcome?).
- Measure the precursor, not only the failure. Prefer leading indicators (e.g. bounded vs climbing context growth) over lagging ones (e.g. a final timeout). The precursor is testable in a short window; the full failure often is not.
- A/B + statistical-N discipline. A single green run is not proof of a behavioral change. Compare a control arm against a treatment arm; require N independent trials and report the pass rate, not a single anecdote.
- Test data fidelity. Validate on real sessions where available, or on faithfully modeled synthetic corpora that preserve the distribution that matters (e.g. heavy-tail file sizes), never on data hand-shaped to pass.
Out of scope (point elsewhere — do not restate)
- Scenario mechanics, the two-layer structure, DTU profile templates / Gitea URL rewrite →
context-intelligence-eval-design. - DTU lifecycle, launch/exec/profiles → the
digital-twin-universeskill. - "DTU-as-default" rationale and the "artifact-as-success" anti-pattern → already authoritative
in
context-intelligence-eval-designandcontext-intelligence:context/context-intelligence-primitives-reference.md. Reference them; do not repeat them.
Method
- Name the user-meaningful outcome the metric must reflect (from
domain-concepts.md), not what a detector happens to emit. - Pick the cheapest axis that discriminates. If a deterministic efficiency signal (token growth, tool-call count) separates good from bad behavior, prefer it over an LLM-judged quality rubric.
- Identify the precursor. Ask: what climbs before the failure? Measure that in a bounded window.
- Design the A/B. Control = current behavior; treatment = the change. Hold everything else identical. Decide N and the pass threshold before running.
- Choose the corpus. Real sessions if available; otherwise synthetic that preserves the load-bearing distribution. State the fidelity assumption explicitly.
Apply (do not restate) the event-semantics principle: when a metric depends on what an event means, consult the authoritative ecosystem expert agents for that event's semantics — the principle is named once in
context-intelligence:context/context-intelligence-strategy.md. Do not restate it here.
More skills from the amplifier-bundle-context-intelligence repository
View all 6 skillsblob-reading
extract fields from blob URIs
Jul 7Data EngineeringFile StorageMicrosoftcontext-intelligence-graph-query
query context intelligence property graphs
Jul 7AgentsData AnalysisGraph AnalysisMicrosoftcontext-intelligence-session-navigation
navigate and extract session data
Jul 7AgentsData AnalysisMicrosoftcontext-intelligence-session-reconstruction
reconstruct local Amplifier session files
Jul 3AmplifierMetadataMicrosoftObservabilityworkflow-pattern-analysis
analyze workflow success and failure patterns
Jul 7Data AnalysisDebuggingMicrosoftWorkflow Automation
More from Microsoft
View publisherrushstack-best-practices
manage Rush monorepos with best practices
rushstack
Apr 6EngineeringLocal DevelopmentMicrosoftProject Management +1azure-ai-agents-persistent-dotnet
build AI agents with Azure .NET SDK
skills
Jul 3.NETAgentsAzureLLMazure-ai-anomalydetector-java
build anomaly detection applications with Java
skills
May 13AnalyticsAzureData AnalysisJava +2azure-ai-contentsafety-java
build content moderation applications with Azure AI
skills
Jul 7AI InfrastructureAzureJavaSecurityazure-ai-contentsafety-py
detect harmful content with Azure AI Content Safety
skills
Jul 18AzureComplianceLLMMicrosoft +2azure-ai-language-conversations-py
implement conversational language understanding with Python
skills
Jul 31AnalyticsAzureLLMPython