Skip to main content

CloudWatch Solutions

27-Point Website & Infrastructure Health Checklist

Work through all 27 before fixing anything. Mark each one Yes, No or Don’t know. A “don’t know” is a finding: not knowing whether backups restore is the same as them not restoring — you just haven’t found out yet.

Then fix in this order: backups, then access and ownership, then updates, then monitoring. Those four remove most of the ways a website turns into an emergency.

Download the PDF

Version 2026-09

Ownership & access

Everything else on this list is fixable. Losing control of the domain, the DNS or the hosting account is not. Start here.

  1. 1

    You control the domain registrar account

    Your answer
    What to check
    You (not an agency, not a former employee) can log in to the registrar. Auto-renew is on, the domain is locked against transfer, and the contact email is one you read.
    Why it matters
    Expired or hijacked domains take the website, email and every login reset with them. Recovery can take weeks and is not guaranteed.
    Red flag
    Red flag: "Our old web guy set that up."
  2. 2

    DNS is hosted somewhere you can log into, and the records are documented

    Your answer
    What to check
    You know which service answers DNS for the domain (registrar, Cloudflare, host) and have the login. Every record has a known purpose; nothing points at a server that no longer exists.
    Why it matters
    Stale records are how a forgotten subdomain becomes a phishing page, and undocumented records are why migrations break email.
    Red flag
    Red flag: Records nobody can explain, or a DNS provider nobody can access.
  3. 3

    The TLS certificate is valid, automatic and complete

    Your answer
    What to check
    HTTPS works on both the bare domain and www, HTTP redirects to HTTPS, the certificate renews itself, and no page loads mixed (insecure) content.
    Why it matters
    An expired certificate is a full outage in every browser. Manual renewal is a reminder somebody eventually misses.
    Red flag
    Red flag: A calendar entry to renew the certificate.
  4. 4

    You know who hosts the site, what you pay, and who holds the login

    Your answer
    What to check
    The hosting company, plan, renewal date and account owner are written down. You can add and remove people’s access yourself.
    Why it matters
    When something breaks at 9pm, the first question is always "where does this run and who can get in".
    Red flag
    Red flag: Hosting billed through a third party you can’t reach directly.
  5. 5

    There is one document listing every account and who has access

    Your answer
    What to check
    Registrar, DNS, hosting, CDN/Cloudflare, email, analytics, premium plugin licences, payment gateways: each with owner, access list and renewal date. Reviewed at least yearly.
    Why it matters
    This is the difference between a two-hour handover and a two-month archaeology project when a provider or employee leaves.
    Red flag
    Red flag: It lives in someone’s head, or in a departed contractor’s inbox.

Backups, recovery & monitoring

Most website emergencies are not caused by attackers. They are caused by an update, a full disk or a bad change, with no way back.

  1. 6

    Backups run automatically and are stored off the server

    Your answer
    What to check
    Files and database are backed up at least daily (more often if the site changes more often), copied to different storage than the server itself, and kept for 30 days or more.
    Why it matters
    A backup on the same disk as the site disappears with the site. A backup from last quarter restores last quarter’s orders.
    Red flag
    Red flag: "The host does backups" — with no idea where they are or how far back they go.
  2. 7

    A restore has actually been tested

    Your answer
    What to check
    Someone has restored a backup to a test location in the last six months, confirmed the site worked, and knows roughly how long it took.
    Why it matters
    Backups that have never been restored are a hope, not a plan. Corrupt archives, missing tables and missing uploads only show up on restore day.
    Red flag
    Red flag: The last restore was never.
  3. 8

    There is a written recovery plan for the likely disasters

    Your answer
    What to check
    Short, plain-language steps for: the server dies, the site is hacked, the domain lapses, the hosting provider disappears. Each names who acts and how long the business can tolerate being down.
    Why it matters
    Decisions made during an outage are worse and slower than decisions made in advance.
    Red flag
    Red flag: "We’d call someone."
  4. 9

    External uptime monitoring alerts a human

    Your answer
    What to check
    A service outside your hosting checks the site every 1–5 minutes, verifies real page content (not just a 200 status), and alerts a phone — not only an inbox.
    Why it matters
    Without it, customers find outages before you do. With it, you often fix them before customers notice.
    Red flag
    Red flag: You learned about the last outage from a customer.
  5. 10

    A staging copy exists for testing changes

    Your answer
    What to check
    There is a non-public copy of the site where updates, new plugins and design changes are tried before they reach production.
    Why it matters
    Every update is a change to running code. Testing it against a copy is the cheapest insurance there is.
    Red flag
    Red flag: Updates are applied straight to the live site and "we watch for problems".
  6. 11

    Updates follow a procedure, on a schedule, with an owner

    Your answer
    What to check
    A named person applies updates at a known cadence, takes a backup first, tests on staging, and checks key pages afterwards. Security updates are handled faster than feature updates.
    Why it matters
    Unpatched software is how most sites get compromised; careless patching is how most sites get broken. A procedure fixes both.
    Red flag
    Red flag: Updates happen "when someone remembers" or never.

Server & platform

Slow sites and sudden outages usually trace back to a handful of unglamorous things: resource headroom, PHP configuration, the database, and caching that isn’t actually working.

  1. 12

    Resource headroom is known and alerted on

    Your answer
    What to check
    You know the server’s CPU, memory and disk usage at a normal peak. Disk stays below ~80%. Someone is alerted before memory or disk run out, not after.
    Why it matters
    A full disk breaks logins, uploads and the database in confusing ways. Memory exhaustion is the most common cause of "the site just went down for no reason".
    Red flag
    Red flag: No graphs, no alerts, or a disk already over 90%.
  2. 13

    PHP is a supported version and PHP-FPM is sized for the traffic

    Your answer
    What to check
    The PHP version still receives security updates. The PHP-FPM pool (worker count, memory per worker) matches the server’s RAM and real traffic, and PHP errors are logged somewhere reviewable.
    Why it matters
    Too few workers and requests queue up under load; too many and the server swaps. Both look like "the site is slow sometimes".
    Red flag
    Red flag: PHP-FPM settings have never been changed from the defaults, or the PHP version is end-of-life.
  3. 14

    The database is healthy and its size is known

    Your answer
    What to check
    MySQL/MariaDB is a supported version, the database size is known and reasonable, the WordPress options table isn’t bloated with autoloaded data, and slow queries can be identified.
    Why it matters
    Database problems are usually gradual (growing tables, transients, logs) and then sudden (timeouts, lock waits, full disk).
    Red flag
    Red flag: Nobody has looked at the database since the site launched.
  4. 15

    Caching is configured and verifiably working

    Your answer
    What to check
    Full-page caching for anonymous visitors, an object cache if the site is dynamic or busy, and sensible browser-cache headers for static files. Verified with response headers, not just a plugin that says "enabled".
    Why it matters
    Caching is the difference between a server that handles a traffic spike and one that falls over. Misconfigured caching also causes stale content and broken carts.
    Red flag
    Red flag: Two caching plugins active, or a cache plugin installed but every page still hits PHP.
  5. 16

    Logs are retained, rotated and reviewable

    Your answer
    What to check
    Web server, PHP and application logs are kept for at least 14–30 days, rotated so they can’t fill the disk, and someone knows where they are.
    Why it matters
    After an incident, logs are how you find out what happened. Without them you are guessing, and you can’t tell if it will happen again.
    Red flag
    Red flag: Logging disabled "to save space", or logs that have filled the disk.

Edge & security

A business website is attacked constantly by automated tools, regardless of how small the business is. The goal is to make the site an unrewarding target and to notice when something gets through.

  1. 17

    Static assets are served through a CDN

    Your answer
    What to check
    Images, scripts and stylesheets are delivered from a CDN edge close to the visitor, with cache rules you understand and can purge.
    Why it matters
    Reduces load on the origin server and makes the site faster for visitors who aren’t near your data center.
    Red flag
    Red flag: Every image on every page is served directly by the origin server.
  2. 18

    Cloudflare (or equivalent) is correctly configured

    Your answer
    What to check
    DNS records for the site are proxied, SSL mode is Full (strict), and the origin server only accepts traffic from the proxy — so protection can’t be bypassed by hitting the server IP directly.
    Why it matters
    A proxy that can be bypassed provides the feeling of protection without the protection.
    Red flag
    Red flag: Cloudflare is "on" but the origin IP is public and answers directly.
  3. 19

    A WAF and rate limiting protect the application

    Your answer
    What to check
    Managed WAF rules are enabled, login and XML-RPC endpoints are rate limited, and obviously malicious traffic is challenged or blocked before it reaches PHP.
    Why it matters
    Most attacks are noisy and automated. A WAF stops the bulk of them at the edge, so the server’s resources go to real visitors.
    Red flag
    Red flag: Thousands of login attempts a day reaching the server unchallenged.
  4. 20

    Admin access is hardened

    Your answer
    What to check
    No account named "admin", strong unique passwords, as few administrator accounts as possible, former staff removed, login attempts limited, and XML-RPC disabled unless something needs it.
    Why it matters
    Credential stuffing and brute force are cheap for attackers. The admin panel is the most valuable door on the site.
    Red flag
    Red flag: Six administrators, two of whom nobody recognises.
  5. 21

    Multi-factor authentication is on for every critical account

    Your answer
    What to check
    MFA is enabled on the registrar, DNS, hosting, Cloudflare, email admin and the website admin — for every person with access.
    Why it matters
    A leaked or reused password is the most likely way any of these accounts is taken over. MFA turns that from a takeover into a failed login.
    Red flag
    Red flag: MFA is "on the important ones", meaning some.
  6. 22

    Malware and file-integrity scanning run on a schedule

    Your answer
    What to check
    A server-level scan (not only a plugin inside WordPress) runs regularly, core files are checked against known-good versions, and someone is told when something changes unexpectedly.
    Why it matters
    Compromised sites are usually discovered late — by Google, a customer or a blacklist. Scanning turns that into an early, quiet notification.
    Red flag
    Red flag: The only scanner is a plugin that the malware can disable.
  7. 23

    Known vulnerabilities in plugins, themes and core are tracked

    Your answer
    What to check
    Installed software is checked against a vulnerability feed, and there is an alert (and a fast-track update path) when something in use gets a published CVE.
    Why it matters
    Attackers automate exploitation within hours of a disclosure. Knowing the same day is the difference between patching and cleaning up.
    Red flag
    Red flag: You’d find out about a plugin vulnerability when the site is defaced.

WordPress & email

The application layer is where most of the day-to-day risk lives: outdated or abandoned software, and email that quietly stops arriving.

  1. 24

    Core, plugins and themes are current, licensed and pruned

    Your answer
    What to check
    Everything is on a supported version, premium licences are valid (so updates arrive), unused plugins and themes are deleted rather than deactivated, and nothing installed has been abandoned by its author.
    Why it matters
    Every installed plugin is code that can be exploited whether or not it’s active. Abandoned plugins never get fixed.
    Red flag
    Red flag: A plugin last updated years ago, or an expired licence that has silently stopped updates.
  2. 25

    Email from the site is authenticated and actually delivered

    Your answer
    What to check
    Transactional email (forms, orders, password resets) is sent via an SMTP or API service rather than the server’s mail function. SPF, DKIM and DMARC are published for the domain. A test message lands in the inbox, not spam.
    Why it matters
    Unauthenticated email from a web server is increasingly rejected outright. The failure is silent: forms "work" and nobody receives them.
    Red flag
    Red flag: Contact form submissions that "sometimes don’t come through".

Performance & accessibility

What the visitor experiences. These two checks are the ones customers and search engines both notice.

  1. 26

    Key pages pass Core Web Vitals

    Your answer
    What to check
    The homepage and top landing pages have acceptable Largest Contentful Paint, interaction responsiveness and layout stability, measured with real-user or lab tools. Time-to-first-byte is reasonable.
    Why it matters
    Slow pages lose visitors before they read anything, and search rankings account for it.
    Red flag
    Red flag: Nobody has measured it, or the measurement is from the launch.
  2. 27

    The site works on a phone and for people using assistive technology

    Your answer
    What to check
    Pages are usable on a small screen without zooming, images are sized for mobile, tap targets are large enough, headings are structured, images have alt text, contrast is adequate, and the site can be navigated by keyboard.
    Why it matters
    Most visitors are on phones; some visitors use screen readers. Both are customers, and accessibility failures carry legal exposure in many places.
    Red flag
    Red flag: It was only ever tested on the designer’s laptop.

Related reading: How to know if your backups actually work · Why monitoring matters before something breaks · What we check before taking over a website

Contact Us for a Confidential Chat

Deliver a better customer experience for your VIPs

Get in Touch