Sechno
Devops

Continuous AI for Accessibility: Turn Axe Results into Actionable Fixes in CI

Practical guide to combining automated accessibility scanners with an AI-assisted CI step that summarizes violations, suggests pragmatic fixes, and posts review-friendly comments. Includes axe + Puppeteer example, CI integration, and an LLM annotation pattern with tradeoffs and mitigation strategies.

SSechno Team 6 min read 74 views
Continuous AI for Accessibility: Turn Axe Results into Actionable Fixes in CI

Why combine automated accessibility checks with AI

Automated scanners (axe-core, pa11y, Lighthouse) catch many programmatic accessibility problems but often return raw rule violations that are noisy for reviewers. Adding a lightweight AI step inside CI can:

  • Summarize violations into human-friendly language
  • Prioritize problems likely to affect users (keyboard, ARIA, alt text)
  • Propose concrete code-level remediation or test suggestions

This pattern improves reviewer productivity and keeps accessibility work in the same dev loop. The examples below use axe-core to collect violations, then call a generic LLM endpoint to translate results into actionable PR comments.

Architecture and tradeoffs

Minimum architecture:

  1. Test runner that crawls pages and runs axe-core (node script).
  2. CI job that uploads the axe report as JSON.
  3. Annotation step that calls an LLM (cloud or self-hosted) to generate summarized comments and suggested fixes, then posts them as PR review comments.

Key tradeoffs:

  • Accuracy vs. noise — LLMs can improve readability, but may hallucinate code changes. Always pair AI suggestions with links to the original rule documentation and an explicit “verify before apply” note.
  • Privacy & compliance — Sending HTML or screenshots to a third-party model can leak sensitive UI content. Use a self-hosted model or redact before sending if that’s a concern.
  • Cost & latency — Invoking an LLM in every commit can add cost and CI time. Run the AI step only when axe finds >= N violations, or run it on PRs targeting main branches.
  • False positives — Automated checks and AI summaries can both misclassify; keep a human-in-the-loop and enable quick suppression metadata in reports.

Step 1: Run axe-core in a headless browser (Node)

This script crawls a single URL, injects axe-core, runs the audit, and writes a JSON report. Install: npm install puppeteer axe-core.

const fs = require('fs');
const puppeteer = require('puppeteer');
 
async function runAxe(url, outPath = 'a11y-report.json') {
  const browser = await puppeteer.launch({ args: ['--no-sandbox'] });
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle2' });
 
  // Inject axe-core from node_modules
  await page.addScriptTag({ path: require.resolve('axe-core/axe.min.js') });
 
  // Run axe in the page context
  const results = await page.evaluate(async () => {
    return await axe.run(document, { runOnly: { type: 'tag', values: ['wcag2a', 'wcag2aa'] } });
  });
 
  await browser.close();
  fs.writeFileSync(outPath, JSON.stringify(results, null, 2));
  console.log('Wrote', outPath);
}
 
if (require.main === module) {
  const url = process.argv[2] || 'http://localhost:3000/';
  runAxe(url).catch(err => { console.error(err); process.exit(1); });
}

Step 2: Add a CI job (GitHub Actions example)

Run the scanner and call the AI annotator only when the report contains violations over a threshold. Keep this step conditional to save cost.

name: a11y-check
 
on:
  pull_request:
    paths:
      - '**/*.html'
      - 'src/**'
 
jobs:
  run-a11y:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Setup Node
        uses: actions/setup-node@v4
        with:
          node-version: '18'
      - name: Install
        run: npm ci
      - name: Run axe scan
        run: node scripts/a11y-scan.js "https://example.org" && echo "A11Y_DONE"
      - name: Upload report
        uses: actions/upload-artifact@v4
        with:
          name: a11y-report
          path: a11y-report.json
      - name: Annotate with AI (conditional)
        if: always()  # make this conditional on # of violations in a real pipeline
        run: python3 scripts/ai_annotate.py a11y-report.json
        env:
          LLM_ENDPOINT: ${{ secrets.LLM_ENDPOINT }}
          LLM_KEY: ${{ secrets.LLM_KEY }}

Step 3: Turn raw violations into review comments using an LLM (Python example)

The annotator reads the axe JSON, formats a compact prompt, and calls a generic LLM endpoint. Keep prompts small: include a short example violation, the selector, and a link to the rule doc. Never send the full page HTML unless you control the model.

import os
import json
import requests
import sys
 
LLM_ENDPOINT = os.environ.get('LLM_ENDPOINT')
LLM_KEY = os.environ.get('LLM_KEY')
 
def load_report(path):
    return json.load(open(path, 'r'))
 
def summarize_violations(report, max_items=10):
    violations = report.get('violations', [])
    summary_items = []
    for v in violations[:max_items]:
        for node in v.get('nodes', [])[:2]:
            summary_items.append({
                'rule': v.get('id'),
                'impact': v.get('impact'),
                'message': v.get('description'),
                'selector': node.get('target', [])[0] if node.get('target') else '',
                'html_snippet': (node.get('html') or '')[:500]
            })
    return summary_items
 
def build_prompt(items):
    prompt = 'You are an accessibility engineer assistant. For each violation, produce:\n'
    prompt += '- a one-sentence explanation of the problem,\n'
    prompt += '- a short, concrete fix suggestion referencing the CSS selector provided,\n'
    prompt += '- a severity note (High/Medium/Low) and a link to the relevant axe rule.\n\n'
    prompt += 'Violations:\n'
    for it in items:
        prompt += f"RULE: {it['rule']} | IMPACT: {it['impact']} | SELECTOR: {it['selector']}\n"
        prompt += f"HTML: {it['html_snippet']}\n---\n"
    return prompt
 
 
def call_llm(prompt):
    # Generic POST to an LLM-compatible endpoint (replace with your provider)
    if not LLM_ENDPOINT or not LLM_KEY:
        print('LLM_ENDPOINT or LLM_KEY missing; skipping AI step.')
        return None
    resp = requests.post(LLM_ENDPOINT, json={'prompt': prompt, 'max_tokens': 600}, headers={'Authorization': f'Bearer {LLM_KEY}'})
    resp.raise_for_status()
    return resp.json().get('text') or resp.text
 
if __name__ == '__main__':
    path = sys.argv[1]
    report = load_report(path)
    items = summarize_violations(report)
    if not items:
        print('No violations to summarize.')
        sys.exit(0)
    prompt = build_prompt(items)
    result = call_llm(prompt)
    if result:
        print('AI summary:\n', result)
        # In CI: post as a PR review comment using the repository's API (not shown here)

Example AI-generated PR comment (format)

Below is a sample of the sort of compact, reviewer-focused comment the annotator should produce. Keep comments short and link back to the raw rule and failing selector so reviewers can reproduce quickly.

Issue: img element missing alt text (axe rule: image-alt)
Severity: High
Selector: main .product-list img:nth-of-type(1)
Explanation: Images without alt text are inaccessible to screen reader users.
Suggested fix: Add meaningful 

Operational recommendations

  • Run the AI annotator only when violations exceed a threshold (e.g., > 5) to reduce noise and cost.
  • Attach the raw axe JSON to the PR so reviewers can inspect the original evidence.
  • Include the rule identifier and a link to the official docs in every AI suggestion to reduce hallucination risk.
  • Maintain a whitelist/ignore list for known false positives and for framework-generated markup you cannot change.
  • Use short prompts and redaction for privacy-sensitive content, or run the annotator against a self-hosted model in high-security environments.

Conclusion

Combining deterministic accessibility scans with an AI layer that summarizes and suggests fixes can make accessibility reviews faster and more actionable. Keep humans in the loop, limit exposure of sensitive UI data, and treat AI suggestions as guidance rather than authoritative changes. With a small CI investment (scan -> annotate -> comment), teams can turn noisy reports into clear, prioritized action items that fit into normal PR workflows.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment