24 WordPress sites, one command per batch: how we automated maintenance with AI (and what broke)

September: 24 WordPress sites, 156 components updated, 52 held back on purpose and 1 real regression among 20 alerts. Method and rules.

Carlos Betancur Gálvez

By Carlos Betancur Gálvez

AI, digital marketing and medical marketing consultant · btodigital

In September 2026 we ran maintenance on the 24 WordPress sites we look after under monthly contracts: 156 components updated, 52 deliberately left alone. Nineteen of those sites went through an automated cycle that Claude Code wrote and operates. One batch of 16 sites took 93 minutes, about 5.8 minutes per site.

This is for agencies, freelancers and in-house IT teams who keep WordPress and Elementor sites alive. Here is the full method and the rules we do not negotiate.

Why did we automate in September of all months?

Because of a scare. On 22 and 23 September, Elementor 4.3.0 installed itself on several sites we manage through automatic updates. Wherever Elementor Pro was still on the 3.x line, the pairing threw a fatal error and the site returned HTTP 500 on every page.

We brought 13 sites back in two rounds by rolling Elementor core down to the last 3.x release that matched their Pro version. Another 7 we fixed before they fell over: same pairing, waiting for the next automatic run. Then we locked down every WordPress install we had access to, so nothing updates unless we decide it should.

That made something uncomfortable obvious. “Manual maintenance, whenever there is time” was not protecting anyone. A site can return a healthy status code and still be broken, and it can break on a Tuesday night with nobody touching it.

What does the ten-step cycle look like?

Every site goes through the same ten steps in the same order. For sites with SSH access, a single command runs the whole batch.

StepWhat it doesWhat it prevents
1. InventoryReads WordPress, PHP, plugin and theme versions, and what is pendingUpdating blind
2. BackupCompressed database, plugins and themes, stored outside public_htmlNo way back, or a backup anyone can download
3. Before screenshotsHome page and key pagesNothing to compare against
4. UpdateEverything pending except what is frozen on purposeBreaking a site with a release we already know is unstable
5. Purge cachesAll of them: object, host, Elementor CSS, cache plugins, CDNSeeing breakage that is only cached, or missing real breakage
6. VerifyHTTP, PHP errors, active plugins before and afterSigning off a site that does not load
7. After screenshotsSame pages at 390, 820 and 1440 px, compared pixel by pixelA site that returns 200 with a broken layout
8. Security scanFile integrity, admin accounts, server logsAn infection nobody sees
9. PDF reportWhat was updated, what was held back and why, findingsA client who has no idea what was done
10. PublishThe report lands in the client portalLost emails

The portal is the Cerebro IA de btodigital dashboard, where each client sees their maintenance history.

There is also a step zero: if the home page is not responding properly before we start, the script touches nothing. In the 26 September batch that happened on a site whose firewall blocked our connection, and the cycle stopped for that site only.

What does the AI do, and what does a person do?

Claude Code wrote the cycle’s scripts, runs them, reviews flagged visual comparisons, writes each site’s notes and publishes the report. It was also the one that rescued the sites during the Elementor incident.

What it does not do: decide which risks to accept with a client, or install our plugin on sites without SSH. A person does that, along with the application passwords, from the admin panel. And when the cycle flags a possible regression, it stops and waits for review before publishing anything.

What happened in the September cycle?

MetricResult
Sites under monthly contract24
Via SSH, through the automated cycle19
Via our own plugin (no SSH)3
By hand from the admin panel1
On a host that blocks automated access1
Components updated156
Components held back on purpose52
26 September batch16 sites in 93 minutes (~5.8 min per site)
1 October batch3 sites in parallel in 13 minutes
One site’s backup, before and after compression577 MB → 135 MB

The 52 held back is my favourite number in the report, because it is the one that shows judgement:

Why it was held backComponents
Elementor 4.3 rule (core and Pro frozen)33
Paid plugin with an inactive licence16
Real or suspected regression in the new release3

A paid plugin without an active licence “updates” without changing version, and if you update the free half without the paid half, the form can break. Better to leave it alone and say so in the report.

Across the SSH sites’ logs, the list of active plugins before and after updating was identical. We check it on every run because, as you will see below, once it was not.

How does the machine know something broke visually?

It compares screenshots. If more than 2% of pixels changed, it flags a possible regression; between 0.5% and 2%, a minor change. In the 26 September batch it compared 123 views (each page on phone, tablet and desktop). It flagged 20 as possible regressions. Only one was real.

The real one: an optimisation plugin, in its new release, blew up one site’s SVG icons. Icons designed for 82 pixels came out huge. We rolled that plugin back on that site and kept it held.

The other 19 were predictable noise:

False positive causeWhy the screenshot changes
Pop-upsThey appear in one capture and not the other
Carousels and slidersEach capture lands on a different slide
Background videoA different frame
Animated textA different word in the headline
Cold cacheThe first capture comes out blank
Host firewallThe capture shows the security check page, not the site

Is 20 alerts for one real failure too many? Not to me. Reviewing 20 image pairs takes minutes. Missing the real one would have meant the client finding giant icons before we did.

What did the security scan find?

We added it in October. On a high-traffic B2B site, the server logs showed 20.5 million requests in 30 days, including 67,241 code injection attempts and 7,522 attempts to enumerate the site’s users. None got in, and the host’s firewall stops most attacks before they ever reach those logs.

The value is not in the scary number but in what shows up next to it:

  • On another site, a paid SSL certificate expiring within weeks that, unlike free ones, does not renew itself. Someone has to obtain the new one and install it by hand.
  • On that same site, two privacy policy PDFs from the previous website that returned 404 and still got 150 visits in a week.

Neither shows up in an “is the site up?” check.

Which rules are non-negotiable?

  1. Automatic updates always off, in three layers: a must-use plugin, a constant in wp-config.php, and the host’s own control panel switch, which you cannot turn off from the command line.
  2. Never update Elementor core unless Pro moves to the same release line. And right now, not even then: core and Pro stay frozen until the 4.3 line proves stable. After a patch release we tested 4 sites and one lost the centring and containers in its header.
  3. If it looks unstyled, think cache first. And verify with the exact URL. A ?nocache= parameter skips the host’s page cache and makes you think it is fixed while it is still broken for everyone else.
  4. A 0.00% difference across every view is suspicious. Everything coming out at exactly zero is not good news; it usually means the comparison is not measuring.
  5. One dedicated SSH key per site, used only for maintenance. If one is compromised, you revoke that one and nothing else.
  6. Backups live outside the public folder. If a browser can download it, it is not a backup. It is a leak.

How many hours does it save?

I do not know, and I would rather say so. We never measured how long the manual method took per site, so any savings figure would be invented.

What we do have is machine time: 93 minutes for 16 sites in one batch, 13 minutes for 3 in parallel. What is missing is human time. In the October cycle we are measuring it like this:

  • Start and end time for each site, which the batch log already records.
  • Minutes of human review per site: false positives, notes and report sign-off.
  • One site done entirely by hand, end to end, as a baseline.

That will give us an honest comparison. If it comes out worse than I expect, I will publish that too.

How do you apply this in your agency or company?

You do not need AI to start. You need order. This is the list in the order I would do it:

  • Turn off automatic updates in all three layers: must-use plugin, wp-config.php and host panel.
  • Write down which components are frozen and why. If you use Elementor, check core and Pro together.
  • Take a compressed backup before every update, outside public_html, and test that it cannot be downloaded.
  • Capture before and after screenshots on phone, tablet and desktop.
  • Purge every cache, not just the host’s.
  • Compare the list of active plugins before and after.
  • Distrust 0.00%.
  • Read the server logs once a month: paid certificates, 404s that still get traffic, intrusion attempts.
  • Deliver a report that lists what was held back and why, not only what was updated.

Once the method is clear, AI speeds it up. Before that, it only speeds up the mess. It is the same reason so many AI projects fail in small businesses. And if you want to look at speed after maintenance, Core Web Vitals and technical SEO covers the order I follow.

Frequently asked questions

Can you automate WordPress maintenance without breaking sites? You can cut the risk a lot, not remove it. The point is not to update faster but to verify afterwards: HTTP, PHP errors, active plugins and visual comparison. In one September batch, out of 123 views compared, one showed a real regression and we rolled it back before the client saw it.

Should I turn off WordPress automatic updates? On sites with Elementor and paid plugins, yes. On 22 and 23 September an automatic Elementor update left sites returning error 500. Updating in a controlled window, with a backup and verification, is safer than letting the site update itself.

Why not move Elementor to the latest version? Because core and Pro have to sit on the same release line, and even then the 4.3 line showed visual regressions on some sites. For now we keep both frozen and explain it in every report.

What should I do if my site looks unstyled after an update? Before changing anything, purge every cache and check the exact URL, without parameters like ?nocache=. In several September cases it was the host’s page cache. If it is still broken after purging, restore the backup.

If you want to set up something like this for your team, it is the kind of work I do as an AI consultant working with Claude.

Share
Related posts