Why continuous accessibility + AI?
Automated accessibility tooling (axe, Lighthouse, pa11y) finds technical issues, but raw reports are noisy and often hard to act on during code review. Adding a model-in-the-loop to summarize, triage, and suggest targeted fixes turns machine output into developer-ready guidance. This post shows an architecture and end-to-end examples you can drop into a CI pipeline so accessibility becomes continuous, fast, and actionable.
Goals for the workflow
- Detect accessibility regressions automatically in CI.
- Summarize and prioritize findings into a short, actionable PR comment.
- Attach precise citations (file/selector/line where possible) so fixes are small and local.
- Keep CI fast and avoid noisy repeat comments.
High-level architecture
- Test runner (axe-core or Lighthouse) produces structured results (JSON).
- Lightweight processor normalizes and deduplicates failures.
- Model summarizer condenses findings into prioritized guidance and suggested code diffs or snippets.
- Bot posts a concise comment on the PR, linking to full report artifact.
Tradeoffs to consider
- Speed vs depth — running full Lighthouse can be slow; prefer targeted axe checks for PRs, and schedule full audits nightly.
- Model cost & latency — keep prompts small by pre-processing and limiting examples to top N failures.
- Hallucination risk — models can invent fixes. Always include the original failing selector and raw rule name so reviewers can validate suggestions.
- Flakiness — make tests deterministic (fixed viewport, mocked network) to reduce noisy failures.
Example: GitHub Actions + axe + AI summarizer
Below is a minimal CI flow: run axe tests with a headless browser, save JSON, run a small PHP summarizer that calls a model endpoint and posts a PR comment. The YAML demonstrates the pipeline; the PHP script shows how to post structured content to a model and to GitHub.
# .github/workflows/a11y.yml
name: accessibility
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
run-a11y:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: '20'
- name: Install axe-cli
run: |
npm ci
npm install -g axe-core@latest axe-cli
- name: Run axe against preview (example uses a playwright / test script)
run: |
# Example: run a script that launches the app and runs axe -> outputs axe-results.json
npm run test:axe -- --output=axe-results.json
- name: Upload artifact
uses: actions/upload-artifact@v4
with:
name: axe-results
path: axe-results.json
- name: Summarize and post comment (PHP bot)
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
MODEL_API_KEY: ${{ secrets.MODEL_API_KEY }}
MODEL_ENDPOINT: ${{ secrets.MODEL_ENDPOINT }}
GITHUB_API_URL: https://api.github.com
run: |
php tools/a11y_summary_bot.php axe-results.json ${{ github.event.pull_request.number }}Note: the test runner step depends on your app setup. Use a headless browser (Playwright, Puppeteer) to reach the PR preview URL or a test-server URL and run axe on relevant pages/components.
PHP summarizer + poster (practical template)
The following PHP script shows a pattern: read axe JSON, normalize top failures, call a model API to summarize and suggest fixes, then post a short PR comment. Replace MODEL_ENDPOINT and request payload with your model provider's contract.
<?php
// tools/a11y_summary_bot.php
// Usage: php a11y_summary_bot.php axe-results.json 123
$argvCount = count($argv);
if ($argvCount < 3) {
echo "Usage: php a11y_summary_bot.php <axe-json> <pr-number>\n";
exit(1);
}
$axeFile = $argv[1];
$prNumber = $argv[2];
$githubToken = getenv('GITHUB_TOKEN');
$modelApiKey = getenv('MODEL_API_KEY');
$modelEndpoint = getenv('MODEL_ENDPOINT');
$repo = getenv('GITHUB_REPOSITORY');
$json = file_get_contents($axeFile);
if ($json === false) {
echo "Cannot read axe JSON: $axeFile\n";
exit(1);
}
$data = json_decode($json, true);
if ($data === null) {
echo "Invalid JSON in $axeFile\n";
exit(1);
}
// Normalize failures: pick top 10 unique rule+selector combos
$issues = [];
foreach (($data['results']['violations'] ?? []) as $violation) {
$rule = $violation['id'] ?? $violation['rule'] ?? 'unknown-rule';
foreach ($violation['nodes'] as $node) {
$selector = $node['target'][0] ?? implode(', ', $node['target'] ?? []);
$key = $rule . '||' . $selector;
if (!isset($issues[$key])) {
$issues[$key] = [
'rule' => $rule,
'impact' => $violation['impact'] ?? 'unknown',
'selector' => $selector,
'html' => substr($node['html'] ?? '', 0, 500),
];
}
}
}
$top = array_slice(array_values($issues), 0, 10);
if (empty($top)) {
// No violations - post small green comment
postComment($repo, $prNumber, $githubToken, "✅ Accessibility checks: no violations detected by axe.");
echo "No issues. Posted success comment.\n";
exit(0);
}
// Build a compact prompt for the summarizer model
$promptParts = [
"You are an accessibility engineer assistant. For each item, return a one-line summary and a 2-3 line concrete fix suggestion. Include the rule id and failing selector.\n",
"Context: Production web app, limited changes preferred, prioritize fixes with highest impact.\n",
"Items:\n",
];
foreach ($top as $item) {
$promptParts[] = "- Rule: {$item['rule']}; Impact: {$item['impact']}; Selector: {$item['selector']}; HTML snippet: " . preg_replace('/\s+/', ' ', $item['html']) . "\n";
}
$prompt = implode('\n', $promptParts);
// Call model API (generic example using curl)
$payload = [
'prompt' => $prompt,
'max_tokens' => 600,
'temperature' => 0.2,
];
$ch = curl_init($modelEndpoint);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
'Content-Type: application/json',
'Authorization: Bearer ' . $modelApiKey,
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($payload));
$response = curl_exec($ch);
$err = curl_error($ch);
curl_close($ch);
if ($response === false) {
echo "Model API error: $err\n";
exit(1);
}
// Parse model response (adjust for your provider's response shape)
$respData = json_decode($response, true);
$summaryText = $respData['summary'] ?? ($respData['choices'][0]['text'] ?? null) ?? $response;
// Build concise PR comment
$comment = "### Accessibility quick report\n";
$comment .= "Found " . count($top) . " unique issues (showing up to 10).\n\n";
$comment .= "**AI summary & suggestions (validate before applying)**:\n\n";
$comment .= "> " . str_replace("\n", "\n> ", trim($summaryText)) . "\n\n";
$comment .= "Full axe report: attached artifact.\n";
$comment .= "\n*Tip: verify selectors and test fix locally before committing.*";
postComment($repo, $prNumber, $githubToken, $comment);
echo "Posted a11y comment to PR #$prNumber\n";
// Helper to post comment via GitHub REST
function postComment($repo, $prNumber, $token, $body) {
$url = "https://api.github.com/repos/" . $repo . "/issues/" . intval($prNumber) . "/comments";
$data = json_encode(['body' => $body]);
$ch = curl_init($url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
'Content-Type: application/json',
'Authorization: token ' . $token,
'User-Agent: a11y-bot'
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, $data);
$res = curl_exec($ch);
$err = curl_error($ch);
curl_close($ch);
if ($res === false) {
echo "GitHub API error: $err\n";
return false;
}
return true;
}Implementation tips & best practices
- Keep the model prompt focused. Send only normalized failures and a short context. That reduces cost and tail latency.
- Always include raw evidence. Post rule id, selector, and a short HTML snippet alongside any AI suggestion so reviewers can validate.
- Debounce comments. If multiple runs generate similar summaries, update the existing bot comment instead of creating new ones to avoid noise.
- Nightly full audits. Run deeper, slower audits (Lighthouse full report) on a schedule and attach the full report artifact for accessibility owners.
- Human-in-the-loop gating. Treat AI suggestions as draft code — require reviewer approval before applying suggested fixes programmatically.
Common pitfalls
- False confidence: Models may produce confident-sounding but incorrect fixes. Mitigate by surfacing rule ids and selectors.
- Flaky selectors: Auto-suggesting changes to dynamic selectors can break tests. Prefer guidance that references role/label attributes rather than brittle class names.
- Over-automation: Auto-committing fixes without review can introduce regressions; use branches and PRs for suggested changes.
Conclusion
Combining structured accessibility tests with a lightweight model summarizer turns noisy audit output into developer-friendly remediation instructions. The pattern above is intentionally minimal: keep tests fast and deterministic in PRs, attach raw evidence to reduce hallucination risk, and use models to shorten the feedback loop rather than replace human judgement. Start with axe & a short-model prompt, then expand to nightly Lighthouse runs and richer remediation workflows as confidence grows.
Further reading
- GitHub: Continuous AI for accessibility — design inspiration for model-in-the-loop accessibility workflows.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment