Sechno
Devops

Continuous Accessibility Testing with AI: Build a Feedback-to-Fix Pipeline for Web Apps

Step-by-step guide to integrate automated accessibility audits with an LLM-powered reviewer in CI — actionable examples, tradeoffs, and a GitHub Actions blueprint to turn audit results into PR comments and suggested fixes.

SSechno Team 6 min read 166 views
Continuous Accessibility Testing with AI: Build a Feedback-to-Fix Pipeline for Web Apps

Why continuous AI + accessibility matters

Automated accessibility checks (axe, Lighthouse) find many problems, but translating raw results into developer-friendly remediation steps takes time. Combining those tools with an AI reviewer in CI can turn noisy test output into concise PR comments, suggested code changes, and a persistent knowledge store. This approach follows the direction outlined by GitHub's work on continuous AI for accessibility (GitHub Engineering).

Goals of this pipeline

  • Run deterministic accessibility audits on every PR and main branch build.
  • Summarize and prioritize findings for developers.
  • Generate suggested code fixes or markup improvements where safe.
  • Store issues and AI feedback to reduce repetition and improve signal over time.

Architecture overview

  1. Headless browser audit (Playwright + axe-core or Lighthouse).
  2. Aggregator service that normalizes findings into a compact JSON schema.
  3. LLM-based reviewer that accepts the normalized issues and returns prioritized suggestions.
  4. CI integration (GitHub Actions) to run audits and post review comments on PRs.
  5. Persistent store (database) to collect issues, AI responses, and manual overrides for continuous improvement.

Step-by-step implementation

1) Run axe-core in CI using Playwright

Install Playwright and axe-core as dev dependencies. The snippet below injects axe-core into a page and runs an accessibility scan; replace http://localhost:3000 with your test URL.

const { chromium } = require('playwright');
const fs = require('fs');
(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('http://localhost:3000');
 
  // Inject axe-core from node_modules
  await page.addScriptTag({ path: require.resolve('axe-core/axe.min.js') });
 
  // Run axe in page context
  const results = await page.evaluate(async () => {
    return await axe.run();
  });
 
  fs.writeFileSync('axe-results.json', JSON.stringify(results, null, 2));
  console.log('Saved axe-results.json with', results.violations.length, 'violations');
  await browser.close();
})();

2) Normalize results to a compact schema

Map axe's result object to a smaller shape for the AI prompt and storage. Keep only necessary fields: id, impact, nodes (selector + html snippet), and help text.

function normalizeAxe(results) {
  return results.violations.map(v => ({
    id: v.id,
    impact: v.impact,
    help: v.help,
    helpUrl: v.helpUrl,
    nodes: v.nodes.map(n => ({
      target: n.target.slice(0, 3), // first few selectors
      html: n.html.replace(/\s+/g, ' ').slice(0, 300)
    }))
  }));
}
 
// usage: const normalized = normalizeAxe(require('./axe-results.json'));

3) Call an LLM reviewer to convert issues into actionable suggestions

Send a concise prompt with the normalized findings and ask for prioritized, low-risk suggestions with code-like examples. Keep the prompt deterministic: include explicit instructions (e.g., prefer semantic HTML fixes, avoid guessing CSS changes that could break layout).

const fetch = require('node-fetch');
 
async function aiReview(normalizedIssues) {
  const prompt = `You are an accessibility reviewer. For each issue provide: priority (high|medium|low), a 1-2 sentence reason, and a suggested code or markup fix. Do not hallucinate project-specific internals.`;
 
  const body = {
    prompt: prompt + '\n\n' + JSON.stringify(normalizedIssues, null, 2),
    max_tokens: 800
  };
 
  const res = await fetch(process.env.LLM_API_URL, {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.LLM_API_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify(body)
  });
 
  if (!res.ok) throw new Error('LLM request failed');
  return await res.json();
}
 
// usage: const review = await aiReview(normalized);

4) Post actionable comments to the PR

Use the GitHub REST API to create review comments. For each suggestion, attach the selector or a small HTML snippet and a short suggested diff. Don't auto-merge; leave the final decision to maintainers.

const fetch = require('node-fetch');
 
async function postPRComments(owner, repo, pull_number, suggestions) {
  const url = `https://api.github.com/repos/${owner}/${repo}/pulls/${pull_number}/reviews`;
  const body = {
    event: 'COMMENT',
    body: suggestions.map(s => `**${s.priority}**: ${s.reason}\n\nSuggested change:\n\n\`\`\`html\n${s.suggested_fix}\n\`\`\``).join('\n---\n')
  };
 
  const res = await fetch(url, {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.GITHUB_TOKEN}`,
      'Accept': 'application/vnd.github+json'
    },
    body: JSON.stringify(body)
  });
  if (!res.ok) throw new Error('Failed to post review');
  return res.json();
}

5) Persist issues and AI responses for continuous learning

Store original findings, AI suggestions, and the developer outcome (applied, ignored, partially applied). Over time you can filter duplicate issues and prefer previously-approved suggestions.

const { Pool } = require('pg');
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
 
async function saveIssue(issue, aiResponse, pr) {
  await pool.query(
    'INSERT INTO accessibility_issues(id, impact, help, nodes, ai_suggestion, pr_number) VALUES($1,$2,$3,$4,$5,$6)',
    [issue.id, issue.impact, issue.help, JSON.stringify(issue.nodes), JSON.stringify(aiResponse), pr]
  );
}

GitHub Actions pipeline (CI) blueprint

Run audits on PRs and main; upload results as an artifact, call your aggregator service, then post back comments. The example below is a compact blueprint to run Playwright tests and save results.

name: accessibility-check
 
on:
  pull_request:
    types: [opened, synchronize, reopened]
  push:
    branches: [main]
 
jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Install dependencies
        run: npm ci
      - name: Start app
        run: npm run start & sleep 5
      - name: Run Playwright accessibility audit
        env:
          CI: true
        run: node ./scripts/run-axe.js
      - name: Upload results
        uses: actions/upload-artifact@v4
        with:
          name: axe-results
          path: axe-results.json
      - name: Call aggregator & AI reviewer
        env:
          LLM_API_URL: ${{ secrets.LLM_API_URL }}
          LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: node ./scripts/ci-annotate.js

Tradeoffs and considerations

  • AI hallucinations: LLMs can suggest fixes that assume internal project structure. Mitigate by constraining prompts, returning only concrete markup changes, and requiring human approval before merges.
  • Privacy and PII: Audit results may contain user data. Avoid sending production content to third-party LLMs or sanitize inputs first.
  • Cost and rate limits: Large repositories and many PRs can drive API costs. Cache common prompts, batch issues per PR, and set thresholds for when to call the LLM (e.g., only high-impact failures).
  • False positives: Automated tools surface many issues—prioritize high-impact items and provide easy ways for developers to mark false positives so the system learns.
  • CI runtime: Headless audits add time. Run full audits on main and quicker audits for PRs (target core pages/components).

Quick checklist before enabling auto-comments

  1. Decide whether AI suggestions are advisory or gating.
  2. Establish prompt templates and allowed-change patterns.
  3. Sanitize inputs sent to external APIs.
  4. Implement human-in-the-loop for the first weeks and collect feedback.
  5. Store outcomes to drive future automation and reduce noise.

Conclusion

Pairing deterministic accessibility tooling with an AI reviewer in CI turns raw findings into developer-ready suggestions and reduces friction for fixing issues. Start small: run axe in CI, normalize results, add an AI reviewer that produces non-destructive suggestions, and iterate based on developer feedback. Over time, a stored corpus of issues and approved fixes will make the system progressively more useful and less noisy.

Further reading: GitHub Engineering: Continuous AI for accessibility, axe-core, Lighthouse.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment