AWS Labs logo

Skill

url-analysis

analyze suspicious URLs in isolated browser sessions

Covers Security Triage Code Analysis Debugging Browser Automation

Description

Analyze a suspicious URL by visiting it in an isolated AgentCore Browser session. Captures screenshots, DOM, network requests, redirects, and extracted IOCs. Use for phishing triage, suspicious-link investigation, and malicious-site fingerprinting.

SKILL.md

url-analysis skill

What this skill does

Analyze a suspicious URL using an isolated AgentCore Browser session, produce a structured forensic report with verdict + confidence + IOCs + recommended actions.

The browser session runs in AWS-managed infrastructure, never in our VPC. Evidence is captured and synthesized into a deterministic verdict via verdict.py.

Your job as the executing agent

Given a URL to analyze:

1. Pre-flight validation

from denylist import DenylistConfig, check_url, scrub_url_credentials

safe_url = scrub_url_credentials(url)
result = check_url(url)
if not result.allowed:
    # Return immediately with status "refused" and result.reason
    # Do NOT create a browser session
    pass

2. Write and execute an orchestration script

Write a Python script that:

  • Opens an AgentCore Browser session (see agentcore-browser-contract.md)
  • Navigates to the URL
  • Captures a screenshot
  • Extracts visible text (via screenshot + your interpretation, or via CDP)
  • Detects any auto-downloads or forms
  • Populates an Evidence object (see evidence_schema.py)
  • Stops the browser session in a finally block

Key rules for the orchestration script:

  • Language: Python 3.11+
  • Use boto3.client('bedrock-agentcore', region_name='us-east-1')
  • Follow agentcore-browser-contract.md for exact API shapes
  • The API provides OS-level actions (mouseClick, keyType, screenshot) NOT high-level browser automation (no navigate, no evaluate, no getHar)
  • To navigate: type the URL into the browser address bar or use Playwright via CDP
  • Screenshots return base64-encoded PNG data
  • Always call stop_browser_session in a finally block
  • Save the script to /tmp/run-artifacts/{run_id}/orchestration.py

Two approaches to browser interaction. Default to CDP. Only fall back to InvokeBrowser if CDP fails at runtime.

  1. Playwright via CDP WebSocket — DEFAULT. USE THIS FIRST.
    Do NOT reject this path because "CDP requires SigV4." The bedrock-agentcore SDK handles SigV4 for you. The one-line idiom is:
    from bedrock_agentcore.tools.browser_client import BrowserClient
    from playwright.sync_api import sync_playwright
    
    bc = BrowserClient(region="us-east-1")
    session_id = bc.start()                     # start_browser_session
    ws_url, headers = bc.generate_ws_headers()  # SigV4-signed, ready to use
    
    with sync_playwright() as p:
        browser = p.chromium.connect_over_cdp(ws_url, headers=headers)
        page = browser.contexts[0].pages[0] if browser.contexts else browser.new_context().new_page()
        page.goto(url, wait_until="networkidle", timeout=30000)
        screenshot_bytes = page.screenshot(full_page=True)
        text = page.inner_text("body")
        # ... extract forms, redirects, etc.
    
    bc.stop()
    

    This gives full DOM access, network interception, form detection, and download events. See examples/001-basic-clean.py for the complete template and examples/004/005/006 for scenario variants.
  2. InvokeBrowser OS actions — FALLBACK ONLY.
    Use only when CDP raises a runtime error you can't work around (e.g., a specific site breaks Playwright, or you need OS-level keyboard input for a native dialog). Lower-level, no DOM access, no form detection, no network interception. Screenshots from InvokeBrowser are full-OS desktop PNGs — resize them before passing to Claude (see Section 9).

3. Enrichment (parallel with browser work if possible)

from enrichment import run_enrichment

enrichment_result = run_enrichment(url, region="us-east-1")

This calls WHOIS, passive DNS, cert transparency, VT, URLhaus, MISP. Each source degrades gracefully if unavailable.

4. Populate Evidence

from evidence_schema import Evidence, ScreenshotCapture, RedirectHop, DetectedForm

evidence = Evidence(
    target_url=url,
    final_url=final_url_after_redirects,
    http_status=200,
    page_title=title,
    screenshots=[ScreenshotCapture(...)],
    visible_text=extracted_text,
    forms=[DetectedForm(...)],
    auto_downloads=[...],
    enrichment={
        "whois": enrichment_result.whois,
        "passive_dns": enrichment_result.passive_dns,
        "cert_transparency": enrichment_result.cert_transparency,
        "virustotal": enrichment_result.virustotal,
        "urlhaus": enrichment_result.urlhaus,
        "misp": enrichment_result.misp,
    },
    run_started_at=start_iso,
    run_completed_at=end_iso,
)

5. Verdict (deterministic - do NOT modify)

from verdict import synthesize_verdict

browser_evidence_dict = evidence.to_browser_evidence_dict()
verdict = synthesize_verdict(
    url=url,
    domain=domain,
    browser_evidence=browser_evidence_dict,
    enrichment=evidence.enrichment,
)

6. Report

from report import render_markdown_report, render_json_report

findings = {
    "url": safe_url,
    "final_url": evidence.final_url,
    "redirect_chain": [r.to_url for r in evidence.redirects],
    "http_status": evidence.http_status,
    "page_title": evidence.page_title,
    "screenshots": [],  # S3 URIs after upload
    "forms_detected": browser_evidence_dict["forms_detected"],
    "auto_downloads": browser_evidence_dict["auto_downloads"],
    "enrichment": evidence.enrichment,
    "iocs": extracted_iocs,
}

md_report = render_markdown_report(safe_url, findings, verdict.to_dict(), duration)

7. Cleanup

Always call stop_browser_session in a finally block. If the session is already terminated, the API returns without error (ResourceNotFoundException is safe to ignore).

8. Screenshot handling (MANDATORY — do not skip)

Browser screenshots at the default viewport (1456×819, full_page=True) can be several MB. Bedrock rejects over-size images with API Error: 400 Could not process image and the whole run dies. Resize before showing to Claude OR keep the screenshot on disk and reason from text evidence.

Before opening a screenshot for visual reasoning, always resize it:

from url_analysis.evidence_store import shrink_for_claude

resized_bytes = shrink_for_claude(screenshot_bytes, max_side=1024)
with open("/tmp/url1_screenshot.png", "wb") as f:
    f.write(resized_bytes)

shrink_for_claude downscales the longest side to max_side pixels and re-encodes as PNG. It's a no-op if the image is already small. Full-resolution bytes stay in the Evidence envelope (uploaded to S3 when the bucket is configured); the on-disk copy is only for Claude's visual input.

If Pillow/PIL is unavailable in the runtime, skip the screenshot read entirely — page.title() + page.inner_text("body") + detected forms give Claude enough to reason from without the image. A missing image must NEVER crash the run.

Example orchestration scripts

See examples/ for reference scripts covering the common scenarios:

#FileScenarioEvidence surface exercised
001001-basic-clean.pyClean URL baselinenavigation, screenshot, forms, text
002002-broken-tls.pyTLS errors (expired, mismatch)graceful degradation, partial evidence
003003-malware-delivery.pyDirect-file delivery (.sh, .dll)page.on("download", ...), SHA-256 without persisting payload
004004-phishing-form.pyCredential harvest / brand-impersonation formspage.evaluate() form enumeration, detached-input detection, brand-host mismatch signals
005005-redirect-chain.pyLink shorteners, cloaking, exploit-kit hopspage.on("response") + page.on("framenavigated")RedirectHop[], TLD-drift + registered-domain-fanout signals
006006-cloudflare-interstitial.pyVendor block pages (Cloudflare / Google SB / SmartScreen)interstitial signature detection, Ray ID extraction, status=partial, do not bypass

Pick the closest match to the URL's signal profile. You can combine patterns — a phishing URL that also uses redirects wants forms from 004 + hop tracking from 005 + the status=partial pattern from 006 if it gets intercepted.

Use these as starting points, not as gospel. The API may drift; if the contract seems wrong, try small experiments and document the real shape in a comment.

Outputs

Stage envelope (JSON):

{
  "artifact_id": "<ARTIFACT_ID>",
  "stage": "url-analysis",
  "stage_name": "url-analysis",
  "timestamp": "<ISO8601 UTC>",
  "status": "ok | partial | failed | refused",
  "duration_seconds": 42,
  "findings": { ... },
  "verdict": {
    "severity": "clean | suspicious | malicious",
    "confidence": 85,
    "category": "phishing | malware-delivery | c2 | scam | unclassified-risk | false-positive",
    "reasoning": "...",
    "mitre_attack": ["T1566.002"],
    "recommended_actions": ["block domain at proxy"]
  },
  "tool_calls": 8,
  "notes": ""
}

Guardrails

  • Never visit internal URLs. Denylist is enforced before any session creation.
  • Never submit forms. Read-only observation of page content.
  • Never click downloads. Detect auto-downloads but don't interact.
  • Session timeout enforced. Default 300s, configurable per-tenant.
  • Credentials scrubbed. Any URL containing auth tokens is masked before persistence.
  • Explicit session termination. Always call StopBrowserSession in a finally block.

Failure handling

  • URL denylist match: refuse immediately, no session created
  • Session creation fails: retry 3x with backoff, then fail with "browser unavailable"
  • Navigation timeout: terminate session, produce partial report with evidence so far
  • Enrichment source unavailable: degrade gracefully, note missing sources
  • Session cleanup fails: log warning, AWS will auto-clean after timeout

© 2026 YourAI.tools. Every skill from an identity-verified publisher.

Independent catalog. Not affiliated with, endorsed by, or sponsored by Anthropic or any listed publisher. All trademarks belong to their respective owners.