seo-audit
Automated SEO, technical SEO, performance and AI-readiness auditor in Python/Django: a URL goes in, 400+ checks, a weighted score and a PDF report come out.
01 Problem
I kept running the same checks by hand on every site I shipped: redirects, canonical tags, sitemaps, structured data, Core Web Vitals, llms.txt. Online checkers each cover a slice, score by rules nobody can see, and often mark a page as fine when they simply could not check it. I wanted one engine that follows my own written checklist, proves every verdict with facts and produces a report a site owner can act on. It started in June 2026 as a personal tool, with the architecture kept ready to run it as a service later.
02 What I built
A Django + DRF service with Celery workers, PostgreSQL and Redis. An audit is an asynchronous pipeline: crawl the site, fetch raw HTML, render pages in Playwright, run Lighthouse, execute the checks, compute the score, build the report. Heavy work (headless Chromium, Lighthouse, PDF) runs only in a separate heavy queue, so a slow render never blocks the light part of the audit.
- Checks are data: a registry generated from a written checklist, with a CI drift check
- No verdict without facts: no pass and no fail without evidence; unverified checks stay out of the score
- Crawler with robots.txt, rate limits and soft-404 probes; Playwright and Lighthouse in a separate heavy queue
- One-language RU or EN reports rendered to PDF; HTTP API with API keys, anti-SSRF and per-IP limits
03 Checks as data, not code
The source of truth is a checklist in my knowledge base: about 40 sections, from HTTP status codes and redirects to JSON-LD, accessibility, analytics and AI crawlers. Each item has an ID, a priority (P0, P1 or P2) and a type: automatic, semi-automatic or manual. A generator turns the checklist into a checks.yaml registry, and a management command fails the build when the registry and the checklist drift apart.
Adding a check means adding a row and a runner function with a test, without touching the engine. That is how the registry grew to 400+ checks in 36 groups. Every runner is tested against a mock site, and the suite of close to 2,900 tests runs without network access.
04 Scoring that does not lie
Each check weighs by priority: P0 is 10 points, P1 is 3, P2 is 1. A group score is the passed weight divided by the weight of applicable checks. Semi-automatic checks count at half weight until an expert confirms them, and manual checks are listed separately. Checks that do not apply leave the denominator: a site without a shop is not penalized for missing product markup.
The rule I care about most: a check cannot pass without a confirmed fact (an HTTP response, a rendered DOM, a measurement), and it cannot fail without one either. If the engine could not verify something, the verdict is "not checked" and it stays out of the score. The summary shows coverage, the share of applicable checks that got an automatic verdict, and flags it when coverage drops below 60%.
05 Crawler, probes and the heavy queue
The crawler uses httpx and selectolax, respects robots.txt, page limits and a requests-per-second limit, and identifies itself with its own user agent. After the main crawl it sends probes: a non-existent URL at the root and in up to eight site sections, plus the http and www variants of the host. The probes caught a real bug on app-kit.dev: a missing blog post returned 200 with noindex instead of 404. After the fix every probe got an honest 404.
Rendering and Lighthouse run per page in the heavy queue with hard time limits. A page that times out is stored as a fact with an error status and never fails the whole audit. The engine also measures image weight without downloading whole files (HEAD, then a one-byte Range request, then a capped stream) and reads the site’s own JavaScript bundles to see when analytics scripts load.
06 A report people can act on
The client report is HTML rendered to an A4 PDF. Every check has its own text: what is wrong, why it matters, how to fix it and how to verify it by hand when automation cannot. Issues are ordered by the points they cost, with the affected pages and the fact behind each verdict. Lighthouse output is translated into plain language, and a test guard stops internal engine terms from leaking into client text.
A report is in one language, Russian or English, chosen per audit or taken from the site’s html lang. The English dictionaries cover every check and several hundred fact labels, and tests make sure the English report has no Cyrillic and the Russian one has no service English. Each run stores a score snapshot, so the history shows how a site changed between audits.
07 From a CLI tool to a public service
The engine runs as a separate service, not inside the App-Kit platform. Crawling arbitrary URLs is an SSRF surface and an unpredictable load, so it lives on its own virtual machine with Docker Compose: API, a light worker, a heavy worker, PostgreSQL and Redis. The HTTP API takes an audit in fast or full mode, checks an API key, rejects private, loopback and cloud metadata addresses before anything is queued (after DNS resolution too) and limits audits per client IP.
Production found bugs that tests had missed: Lighthouse could not find Chromium in the image, Chromium crashed on the default 64 MB of shared memory, and Celery silently dropped the PDF task because its app was not registered. All three are fixed. On App-Kit the engine powers a free online SEO audit: a fast preview in seconds, then the full audit with live progress and a PDF. The article Free Online SEO Audit: 400-Point Check, Score and PDF explains how it works, and the technical SEO audit checklist covers the checks that matter most.
08 Result
The first fast audit of app-kit.ru scored 70.5 out of 100 in 19.6 seconds and listed 14 concrete failures, among them http without a 301 to https, a legacy Host directive in robots.txt, a missing self-referencing canonical and overly long titles. After the fixes, the full audit of 18 pages gave 93% with no losses on P0 checks, and after the next round the fast audit of app-kit.ru reached 98.5%. A full audit with rendering and Lighthouse takes about four and a half minutes. Since June 2026 the project has grown to 144 commits, and it audits my own sites, this one included.