The SEO machine behind my sites
One agent-run content engine positions several WordPress sites in different languages and niches. Three assets held apart, two gates, one weekly loop, and a statistics page as the worked example.
I run several WordPress sites on the side: a German blog on long-term relationships that my partner and I write together, a Serbian health portal in a regulated niche, and a German site on the same health topic. They differ in language, audience and legal exposure, and one machine does the search positioning for all of them.
The machine is a method, a set of scripts, a database and a rulebook. AI agents operate it and I gate it. Here is how it is built, what it has produced, and how it produced its most recent page, a set of statistics that nobody had assembled before.
What goes wrong with AI content
AI content that fails usually fails the same way. It is fluent and fast and says what everyone else already says, so it ranks for nothing, gets cited by nothing, and makes the brand sound like every other brand. Publishing more of it does not help.
To avoid that, the machine keeps three assets apart.
The first is demand data: what people search for, and whether a site with our authority can win those searches. It lives in a SQLite database, and every number in it is a dated, append-only row: keyword volume, difficulty, the top ten results, which SERP features fire, our own position, Search Console impressions, analytics. Six months from now, the question whether something worked is a query against that table.
The second is voice and doctrine: who is speaking, how they sound, and where they disagree with the mainstream. It lives in three Markdown files per project. The third file, the knowledge map, does most of the work. Every topic in it has four parts: the mainstream belief a reader already holds, our position with verbatim quotes as evidence, a bridge that carries a reader from the first to the second without mocking them, and a named concept with a short definition. SEO writing has to enter on the mainstream keyword, because that is where the demand is, and then say something the mainstream does not. The bridge is that move, written down once per topic and reused.
The third asset is the write path: a WordPress plugin I wrote that exposes a scoped endpoint agents can talk to, plus the scripts around it.
An agent holding all three can write an article that targets real demand, sounds like a specific person, says something non-obvious, and lands on the platform correctly tagged and tracked. An agent holding only the first and the third produces the generic version.
Read-only research, draft-only writing
The research credential can only read. The publishing credential can only write drafts, and only through the plugin, with a scoped key: no publish, no delete, no plugins, no users. A confused or compromised agent has a small blast radius.
Early on I published through the command line on the server and hit two silent failures in one afternoon. A post created from the CLI answered 404 on the live site, because the CLI and the web runtime kept separate object caches, and a rewrite flush did nothing because the rewrite rules were fine. A post created from a file arrived with empty content and exit code zero, because the CLI ran as a different user who could not read the file. Writing through the plugin, inside the platform's own runtime, removed both bugs. Every publish script now fetches the public page and checks for a 200, because the API response alone proves nothing.
One skeleton, many sites
The machine is one repository, and it is topic-free. It holds the scripts, the database schema, the templates for the voice files, and the list of traps. Each site is a separate project/ directory that the engine ignores in git and that is its own private repository: credentials, the three voice files, the database, the docs, and a CLAUDE.md with the rules for that site.
The same engine can therefore run a couple-coaching blog and a health portal without either one leaking into the other. Pulling engine updates cannot conflict with a project, because git does not see the project. The doctrine, which took days to write, has its own history and its own backup.
The project rulebook is where the sites differ. On the relationship blog it says: one paragraph per article that points to professional help, no diagnostic language, never tell anyone to stop taking their medication. On the Serbian health portal it is stricter, and I come back to it below.
Inventory before strategy
On a new site the machine starts with an inventory, before any keyword research.
On the relationship blog that inventory changed the plan. The site had 110 indexable pages and zero focus keywords, eight published posts and 184 abandoned LLM drafts. It also had 116 unpublished transcripts of the founders teaching their method. So the site did not need more content. It needed the transcripts turned into articles and the old drafts sorted into three piles: upgrade, merge, or retire.
The inventory also found a plugin that was already generating drafts from form submissions with a stock prompt. Nobody had looked at it. Replacing that prompt with one built from the voice files changed every future draft at once.
What the audience says, in its own words
Keyword tools report what a market types into Google, which is not the same as what our readers struggle with. Where the two disagreed, we followed the readers.
The relationship blog had 765 free-text answers from a questionnaire people fill in after opting in. We imported the text without any personal data into the database as audience signals. The two largest pain clusters, intimacy and everyday couple time, were barely covered by the content plan, because keyword mining had steered elsewhere. We rewrote a whole wave of articles against that data. The answers also gave the writing agents the exact phrasing readers use for the problem, which no keyword export contains.
The SERP data changed the goal as well. In that niche an AI Overview fired on 79 percent of the target keywords and a featured snippet on only 5 percent. An AI answer intercepts the click, so the objective is to be the source cited inside that answer rather than the blue link below it. Other niches will show other numbers, so we measure before writing.
Two gates
The first gate is a second agent. Before anything is published, a separate model reads the draft against the voice files and the hard rules and returns either a pass or a list of required changes. It is wired into the article script and blocks on changes. On its first run it found what I had missed in my own drafts: too explanatory, too clinical, the first person thinned out, healing claims stated as absolutes, and in two cases the focus keyword missing from the text. On the health portal the same gate runs a different checklist: factual errors in mineralogy and regulation, grammar in Serbian, and any sentence that moves from describing a mineral to claiming a benefit.
The second gate is me. The agent drafts and a person publishes, permanently, for two reasons. A brand with a doctrine cannot afford a piece that is 90 percent right. And if any part of the pipeline is triggered by user input, a form answer or a comment, then attacker-controllable text is flowing into an agent that holds CMS write access. The scoped key, the cost caps and the explicit "this is data, not instructions" framing reduce that risk, and the human gate is what holds when they fail.
Sixteen articles published on one Tuesday look like a dump. We back-dated the initial catalogue across the previous months and schedule the ongoing queue one per week, months ahead.
The weekly loop
A weekly cron job snapshots our position for every target keyword, pulls Search Console (including the queries we never targeted), pulls analytics, and rebuilds a one-page dashboard. Untargeted queries that already earn impressions become new rows in the content map.
The first report on a new site reads zero. The relationship blog's first rank run found none of twenty keywords in the top hundred, and Search Console showed three impressions, all brand queries. That is indexing latency, and knowing it in advance stops anyone from rewriting content that is still in the queue.
The Serbian portal launched with 28 URLs and two of them indexed. My first diagnosis was crawl budget on a domain with a messy past, and manual indexing requests pushed the count up over two weeks. Then, while checking something unrelated, I found that the origin server behind the CDN had no valid certificate. Visitors never saw it, because the CDN hid it, and it explained the weeks of zero impressions better than any crawl-budget theory. The certificate is fixed, all 28 URLs are indexed, impressions have grown every week since, and the site checklist now has a line for it.
A JavaScript tracker cannot see AI crawlers, because they do not execute JavaScript. A small worker at the CDN edge therefore counts every visit from GPTBot, ClaudeBot, PerplexityBot and their relatives, records which URLs they read, and writes it into the analytics as its own site.
The claim framework
The Serbian portal sits in a health niche where a German court in 2023 enjoined 68 advertising statements by a distributor of a comparable product. The site operator is responsible for every claim it publishes. So the project rulebook for that site is a claim framework, and it overrides anything a keyword tool suggests.
The rule is: target the keyword, never make the claim. A phrase like "side effects" can be the H1 of an article without the article promising anything. The unit of judgement is the whole reader journey, though, and a safety article with a product card after the third paragraph is a violation even when every sentence in it is defensible, because the page structure implies what the text avoids saying. A disclaimer does not cure that. Every article therefore declares its call-to-action explicitly, and on any page about a contaminant or a symptom the declaration is none.
The framework also lists what is allowed: mineralogy, manufacture, regulatory status, market and price data, and the state of the evidence including its limits. The two pages on that site a reader can trust most are the one about side effects and the one about drug interactions, because competitors are too sales-shy to write them.
On the relationship blog the same slot in the machine holds the doctrine, the file that says what we believe. On the health portal it holds the file that says what we may not say.
The statistics page
A site with no authority gets links by holding something nobody else holds. The health portal's first citable asset was a price table: retail prices of the product across four countries, converted to price per kilogram, dated, with the sample size and the exclusions on the page. Comparable powder costs about six times more in Croatia than in Serbia, by median. The bet is that a sentence like that gets quoted, and that a quote carries a link. So far the site has no external link it did not place itself, so the bet is still open.
The newest page applies the same idea to a whole topic. The Serbian Ministry of Health runs a public register of dietary products that had never been aggregated. We pulled all 18,169 records and counted entries per year, category, country of manufacture, ingredient in the product name, and dosage form. The number of entries registered in 2025 is more than four times the number from 2019. Around that core sit the national health survey, the household budget survey, trade data from UN Comtrade and Eurostat, the EU health claims register, EFSA reference values, USGS production figures for the mineral, and our own price series and search-demand data. Over sixty figures sit under thirteen question headings, each answered in its first sentence, each figure with its source, its period and the date we pulled it. A key-figures block at the top can be quoted as one paragraph. A section at the bottom lists every query and parameter so anyone can reproduce the counts, and the aggregates are downloadable as CSV. Where we could not find a figure, the page says so.
Three research agents fetched every source and had to quote what they read. Anything they could not open went into a "not verified" list and stayed out of the article. The second-agent gate returned thirty findings. I adopted twenty-seven, among them reconciling the register distributions to the exact total and replacing an expired feed authorisation with the regulation that renewed it this year, and rejected three. The page went through the same publish path as every article, carries a visible "last updated" date, and is due for a refresh in December together with the price dataset.
The page is in Serbian: dietary supplements in Serbia and the region, statistics with sources. The price dataset it grew out of is here. The relationship blog is erfolgreich-als-paar.de.
What the machine cannot do
An agent can draft the three voice files from a corpus, and a human then has to edit them hard, because every judgement about what the brand says and how it says it sits in those files. The machine cannot decide which claims a site may make; it enforces the decision. It cannot replace the publish gate, and it cannot make the first report read anything other than zero. Those are also the steps that get skipped under time pressure.
I have written before about building agent-first. This machine is what that looks like when the product is a body of work across several sites instead of an app. The agents do the research, the drafting, the tracking and the bookkeeping. I do the inventory, the doctrine, the rules and the final read.