Sechno
Web Development

Build a Continuous Accessibility Feedback Loop: From Axe to AI-driven Insights

A practical guide for developers to automate accessibility testing in CI, collect human + AI feedback, and turn reports into actionable insights that improve inclusion over time.

SSechno Team 6 min read 91 views
Build a Continuous Accessibility Feedback Loop: From Axe to AI-driven Insights

Why continuous accessibility matters

Automated accessibility tests are useful, but running a scan once isnt enough. Accessibility regressions creep in with UI changes, content updates, and third-party widgets. A continuous feedback loop—automated scans in CI, developer-facing annotations, human review, and AI-assisted summarization—helps teams catch, prioritize, and learn from accessibility issues without slowing delivery.

What this article shows

  • How to run axe-core accessibility scans in Playwright tests
  • How to annotate CI failures (GitHub Checks) so developers see actionable guidance in pull requests
  • How to store and summarize violations with a minimal AI-driven feedback pipeline
  • Tradeoffs, flakiness mitigation, and privacy considerations

1) Run automated scans with Playwright + axe-core

Use Playwright to visit pages and inject axe-core. Running scans as part of your test suite gives reproducible output (JSON) you can forward to CI.

import { test, expect } from '@playwright/test'
import AxeBuilder from '@axe-core/playwright'
 
test('homepage should have no severe accessibility violations', async ({ page }) => {
  await page.goto('https://example.com')
 
  const results = await new AxeBuilder({ page }).analyze()
 
  // Fail test if there are 'serious' or 'critical' violations
  const blocking = results.violations.filter(v => {
    return ['serious', 'critical'].includes(v.impact)
  })
 
  if (blocking.length > 0) {
    // Serialize results for CI artifact upload
    console.log('AXE_VIOLATIONS_JSON=' + JSON.stringify(results))
    throw new Error(`Accessibility regressions: ${blocking.length} issues`) 
  }
})

How this helps: the test fails on important issues and prints a JSON payload you can collect as a CI artifact for later analysis or human review.

2) Surface violations in pull requests (GitHub Checks API)

Instead of making developers hunt artifacts, post annotations on the pull request. Below is a minimal Node snippet that creates a GitHub check and adds annotations for each violation. Run this script as a step in CI after tests produce the axe JSON.

import { Octokit } from '@octokit/rest'
import fs from 'fs'
 
const token = process.env.GITHUB_TOKEN
const octokit = new Octokit({ auth: token })
 
async function annotate(orbUrl) {
  const repo = process.env.GITHUB_REPOSITORY // owner/repo
  const [owner, repoName] = repo.split('/')
  const head_sha = process.env.GITHUB_SHA
 
  const violations = JSON.parse(fs.readFileSync('axe-results.json', 'utf-8'))
 
  // Create a check run
  const check = await octokit.checks.create({
    owner,
    repo: repoName,
    name: 'accessibility/scan',
    head_sha,
    status: 'completed',
    conclusion: violations.violations.length === 0 ? 'success' : 'neutral'
  })
 
  // Build annotations (GitHub limits: 50 per request, 1000 total size etc.)
  const annotations = violations.violations.flatMap(v =>
    v.nodes.map(node => ({
      path: node.target[0] || 'unknown',
      start_line: 1,
      end_line: 1,
      annotation_level: 'warning',
      message: `${v.id}: ${v.help} -- ${node.html}`.slice(0, 64000)
    }))
  )
 
  // Send annotations in chunks
  const chunkSize = 50
  for (let i = 0; i < annotations.length; i += chunkSize) {
    const chunk = annotations.slice(i, i + chunkSize)
    await octokit.checks.update({
      owner,
      repo: repoName,
      check_run_id: check.data.id,
      output: {
        title: 'Accessibility scan results',
        summary: `${violations.violations.length} violations found`,
        annotations: chunk
      }
    })
  }
}
 
annotate().catch(err => { console.error(err); process.exit(1) })

Notes: use a token with repo:status and checks:write. Respect API limits and trim messages to avoid errors.

3) Collect and summarize reports with a lightweight AI

Store scan results and developer feedback (accepted, ignored, fixed) in a simple store. Periodically run a summarizer to produce prioritized recommendations for the team. Below is a Python example that ingests a saved JSON scan and writes a compact summary record to SQLite for later analysis or embedding.

import json
import sqlite3
from datetime import datetime
 
conn = sqlite3.connect('a11y_reports.db')
conn.execute('''CREATE TABLE IF NOT EXISTS reports (
  id INTEGER PRIMARY KEY AUTOINCREMENT,
  page TEXT,
  rule_id TEXT,
  impact TEXT,
  html_snippet TEXT,
  first_seen TEXT
)''')
 
with open('axe-results.json') as f:
    data = json.load(f)
 
for v in data['violations']:
    for node in v['nodes']:
        conn.execute('''INSERT INTO reports (page, rule_id, impact, html_snippet, first_seen)
                        VALUES (?, ?, ?, ?, ?)''', (
            data.get('url', 'unknown'),
            v['id'],
            v.get('impact', ''),
            node.get('html', '')[:2000],
            datetime.utcnow().isoformat()
        ))
 
conn.commit()
conn.close()

Next steps: use these stored rows to feed an AI summarizer (local or cloud) that groups similar violations, surfaces frequently broken rules, and suggests remediation patterns. You can use vector embeddings for grouping but be mindful of sensitive content and PII.

4) Practical pipeline architecture

  1. CI step: run Playwright + axe-core on key pages/components → produce JSON artifact.
  2. CI step: run annotation script → create GitHub check / PR annotations.
  3. Storage: upload JSON to artifact store and ingest a summary row into your store (SQLite, Postgres, or vector DB).
  4. Human review: QA or accessibility reviewers triage annotated PRs and mark fixes or false positives.
  5. AI summarization (daily/weekly): group issues, suggest fixes, generate education snippets for dev docs.

Example metrics to track

  • Number of violations by impact (critical, serious)
  • Time-to-fix accessibility issues
  • Flaky-failure rate for accessibility tests
  • Percentage of violations confirmed as true positives by reviewers

5) Tradeoffs and practical tips

  • False positives: Axe is conservative but can surface issues that need context. Use rule whitelists or only fail on certain impacts to reduce noise.
  • Flakiness: Network delays, lazy-loaded content, and animations cause intermittent failures. Stabilize pages (wait for network idle, ensure fonts/content loaded) before scanning.
  • Developer friction: Too many PR annotations can be noisy. Aggregate annotations, provide clear remediation hints, and prioritize by impact.
  • Privacy and compliance: Scan outputs may contain user-generated content. Avoid sending sensitive HTML to external AI services unless you have consent and masking in place.
  • Cost: Storing long histories or using hosted embeddings/AI can incur cost—sample or downsample reports if needed.

6) Human-in-the-loop and governance

Automated tools are not a replacement for human judgment. Build a lightweight workflow where: QA verifies critical failures, product owners accept high-priority changes, and accessibility champions provide guidance. Keep a changelog of rule exceptions and why they were accepted.

7) Further automation: continuous learning

Over time you can use the stored reports plus reviewer labels to train or tune an internal classifier that predicts whether a violation is actionable for your app. Options include:

  • Simple heuristics and rule filters (low-cost, transparent)
  • Embedding-based grouping to find recurring HTML patterns (requires vector DB)
  • LLM summarizers to generate fix snippets for common issues (be careful with PII)

Minimal safety pattern

  1. Mask or remove text/content before sending to cloud AI.
  2. Keep human review for final acceptance of automated recommendations.
  3. Maintain an audit log of all automated changes or suggestions.

8) Resources and references

Conclusion

Building a continuous accessibility feedback loop is both practical and high-impact: automate scans in CI, annotate pull requests so developers get immediate context, store and summarize violations, and keep humans in the loop for judgment. Start small—scan a handful of critical pages, annotate PRs, and iterate on filters and reporting. Over time the pipeline becomes a learning system that helps teams ship more inclusive software with predictable effort.

Quick checklist: Add Playwright+axe to CI, produce JSON artifacts, annotate PRs with GitHub Checks, store results, and run periodic AI summaries with strict privacy safeguards.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment