AI-assisted content operations

What I Learned From Automating Keyword Research for My Personal Website

A first-party account of automating keyword research for a personal website—what worked, what failed, and the controls that made the output usable.

Huiyang Xie

On this page
  1. I defined the output before automating the input
  2. The first failure happened before the provider call
  3. Raw data became a stage-boundary requirement
  4. Collection was the easy part; interpretation created the backlog
  5. Diminishing returns needed an explicit stop rule
  6. SERP validation changed query strings into reader problems
  7. Scoring worked only because different topics followed different logic
  8. The useful automation ended at a decision-ready state
  9. What I would automate again
  10. Frequently asked questions
  11. See the wider content system

When I began researching topics for HuiyangXie.ca, I did not want a large spreadsheet simply because keyword tools could produce one. I needed a usable answer to a narrower question: which topics could connect real search behavior with work I could discuss credibly on a personal website?

Automation’s contribution was different from the common “automated keyword research” pitch. It created a controlled research chain that made raw evidence easier to collect, inspect, clean, compare, and hand off. Editorial decisions remained a separate responsibility; no AI-generated calendar determined what the site would publish.

That separation became the central lesson of the experiment. Keyword-research automation is reliable when it preserves the difference between what a provider returned, what the workflow inferred, and what a person later approved for production.

This article covers the research subsystem behind my AI-assisted SEO content pipeline. It describes one first-party experiment, not a universal benchmark or evidence that the resulting articles rank.

I defined the output before automating the input

It is easy to begin with a list of tasks: call a provider, collect suggestions, remove duplicates, cluster the results, and assign scores. That sounds like a workflow, but it does not yet define what a good result looks like.

For my site, the desired output was a decision-ready editorial backlog. Each candidate needed more than a phrase. It needed a topic family, likely intent, portfolio role, evidence path, refresh status, cannibalization boundary, and production state. A candidate could be promising without being ready. It could be ready without being approved. A high score could support a recommendation without granting publication authority.

Those distinctions shaped the automation. The system could prepare evidence and surface conflicts. It could not silently turn a row into an article.

I also kept three forms of evidence separate:

  • Provider evidence showed the suggestions returned for a seed under a recorded locale and run.
  • Public search evidence helped interpret current intent, common result formats, saturation, and gaps.
  • First-party authority evidence established whether I had a defensible experience, process, or decision to contribute.

Combining them too early would have made the output look more certain than it was. Keeping them separate made later scoring and editorial review more honest.

The first failure happened before the provider call

The initial batch stopped on its first intended paid run. Read-only preflight commands had confirmed the provider endpoint, documented schema, price at that time, authentication, and account balance. The paid execution context then failed locally before a provider run was created.

The important action was to stop. I left the credentials and configuration untouched instead of guessing at a workaround or trying several paid requests. The ending balance remained unchanged, and the batch recorded one pre-provider execution failure with zero paid calls.

This was a useful design test. A workflow that only knows how to continue is not reliable. It also needs to recognize when the execution state is ambiguous and preserve enough evidence for a later diagnosis.

The failure created a clear stop rule:

  • no repeat attempt while credential or execution behavior is unclear;
  • no claim that the provider failed when the request never reached it;
  • no paid diagnostic call simply to test broken access;
  • no downstream artifact pretending that missing data had been collected.

These rules were more valuable than a brittle workaround. They protected spend, provenance, and the integrity of every later stage.

Raw data became a stage-boundary requirement

When the provider results were later available, the next problem was not the content of the suggestions. It was the handoff.

Downstream cleaning could not safely depend on chat attachments, terminal output, or task history. Those surfaces can help during a session, but they are not stable data contracts. The complete provider outputs had to be restored to a canonical stage-boundary directory, then checked for presence, readability, valid JSON, non-empty payloads, and correct seed mapping before cleaning resumed.

That correction changed the workflow permanently. Every paid response now needs to be persisted immediately with its run context. A later stage should be able to answer basic questions without reconstructing the past:

  • Which input produced this file?
  • Is the payload complete and readable?
  • What did the provider return before cleaning?
  • What cost and run state were recorded?
  • Has this research already been completed and locked?

The principle is simple: transient context can assist a session, while paid research needs a durable canonical handoff.

This also prevents an expensive failure pattern. If an English audit, Chinese revision, website build, or layout check fails weeks later, that downstream problem must not replay the original provider calls. Research reopens only when the research itself has been invalidated, with any new budget decision handled separately.

Collection was the easy part; interpretation created the backlog

The restored experiment contained ten provider calls and 97 raw occurrences. Normalization produced 87 unique queries. The cleaning pass classified 54 as KEEP, 13 as REVIEW, and 20 as NOISE.

Those numbers describe this run only. They are useful because they show why a suggestion list could not go directly into editorial planning.

Some phrases were clear and relevant. Others were ambiguous strings that needed context. AI-writing seeds also attracted pronunciation and phonics queries because “AI” can be interpreted as letters or a sound rather than artificial intelligence. A rule that accepted every suggestion would have produced a misleading view of audience demand.

The opposite rule—discard everything unusual—would also have failed. One seed returned zero suggestions, yet the underlying multicultural-marketing topic still had a strong first-party authority path. The absence of an autocomplete footprint did not prove the absence of reader value.

I therefore separated mechanical cleanup from strategic preservation. Normalization could standardize text, remove exact duplicates, and attach provenance. Classification could flag obvious noise. A topic with weak provider output still needed a deliberate decision about authority, portfolio value, and evidence readiness.

Research layerWhat automation can prepareWhat still needs interpretation
CollectionRun records, raw payloads, locale, cost, and seed provenanceWhether the source is suitable for the decision
NormalizationConsistent text, exact deduplication, and traceable variantsWhether similar phrases represent the same intent
ClassificationKEEP, REVIEW, and NOISE candidates with reasonsWhether a weak search footprint should override real authority
ValidationComparable result notes and repeated SERP patternsWhat gap the site can credibly fill
PrioritizationScores, dependencies, and conflict flagsWhat should enter the editorial queue

The table is also a warning against collapsing every step into one AI score. Each layer answers a different question. A clean phrase is not necessarily a useful topic. A visible query is not necessarily a good fit. A strong professional topic may deserve preservation even when provider coverage is thin.

Diminishing returns needed an explicit stop rule

More seeds do not always create more insight. In one closely related workflow call, seven suggestions produced four exact overlaps and only three new queries. The small expansion was recorded, but it did not justify continuing through a series of near-duplicate paid seeds.

This is where automation can create false confidence. A system can keep collecting because the endpoint still returns data. The growing row count looks productive even when the marginal information is collapsing.

I used overlap and family coverage as stopping evidence. Completion was defined narrowly: the current families contained enough diversity for validation, and another similar paid call was unlikely to change the editorial decision.

The same logic applied to cost. The ten restored call artifacts reported an aggregate price of US$0.324. The observed balance change was US$0.36, leaving US$0.036 unreconciled. I recorded the difference without inventing an explanation. Cost control includes preserving small discrepancies, not smoothing them away because the total is modest.

SERP validation changed query strings into reader problems

Provider suggestions supplied language patterns, but they did not establish what searchers expected to find. The next stage sampled current public results for representative queries and recorded intent, common formats, repeated angles, gaps, and source context.

In the original cycle, 16 representative queries and 81 public-web result samples supported a 13-topic candidate pool. Those samples did not create a keyword-difficulty score or predict rankings. They helped answer editorial questions: Is the result page mostly definitions, tool lists, workflows, or practitioner guidance? Are several phrasings serving one underlying intent? Is there room for a first-party contribution that does not repeat what is already common?

The current focused refresh for this article showed a similar pattern around keyword-research automation. Many pages emphasize connecting tools, routing records, automatic tagging, alerts, and faster planning. Those capabilities can be useful. The underexplored problem is how to keep the automated output auditable when it contains noise, incomplete coverage, uncertain costs, or an interrupted handoff.

That gap became the article’s purpose: an account of the controls required before automation can support an editorial decision, rather than a product comparison.

Scoring worked only because different topics followed different logic

The candidate pool included Search-led, Authority-led, and Experiment-led topics. Applying one formula to all three would have favored visible search phrases and penalized topics grounded in professional experience or an unfinished first-party experiment.

Search-led topics could weigh validated query intent and result gaps more heavily. Authority-led topics needed a defensible practitioner angle and portfolio role even when autocomplete was weak. Experiment-led topics required real process or outcome evidence and a readiness gate. An experiment could be strategically valuable while remaining ineligible for publication until the evidence matured.

That distinction protected two kinds of honesty:

  • search evidence was not inflated into professional authority;
  • professional authority was not dismissed merely because a provider returned few suggestions.

Cannibalization was evaluated separately. Several AI-writing query variants became one comprehensive article rather than a cluster of near-duplicates. Other topics stayed separate only when they answered different reader questions. The score informed that work; it did not replace it.

The useful automation ended at a decision-ready state

The operating model that emerged has five connected responsibilities:

  1. Collect with controls. Validate the source, scope, locale, price, budget, and run state before a paid request.
  2. Persist before transforming. Save the raw response and provenance where downstream work can read and verify them.
  3. Normalize without erasing meaning. Clean mechanically, preserve source variants, and label uncertainty rather than hiding it.
  4. Interpret with the right evidence. Use SERP context, authority fit, experiment maturity, and cannibalization to understand what a query could become.
  5. Decide through explicit state. Keep backlog, recommendation, production approval, and publication as separate transitions.

This five-part model reflects the smallest structure that made my experiment inspectable and resumable. Another website may need a different shape.

After publication, another evidence layer should enter the system. Google’s documentation distinguishes Search Console search-performance data from behavior measured inside a site, and its Performance report guidance shows how queries and pages can be examined. Those observations can eventually inform backlog and refresh decisions. They still need a meaningful window and careful interpretation; they do not retroactively prove that the original keyword score was correct.

What I would automate again

I would automate the repetitive handling again: controlled collection, raw persistence, normalization, provenance, structured audit outputs, and consistent state records. These tasks benefit from repeatability and are easy to verify.

I would not ask the automation to decide what the website should stand for. Topic selection still depends on professional direction, evidence that can be disclosed, portfolio balance, and the cost of occupying a reader’s attention with one article instead of another.

The most valuable output was research that another stage could read, verify, question, and use without guessing what had happened earlier. The size of the keyword universe was secondary.

Questions

Frequently asked questions

What parts of keyword research can be automated?

Collection, normalization, deduplication, labeling support, artifact generation, and repeatable validation checks can often be automated. Topic interpretation, evidence readiness, cannibalization decisions, budget authority, and publication priority still require an explicit decision model and, in this workflow, owner authority.

Why save raw keyword data before cleaning it?

Raw data preserves what the provider actually returned. It allows later stages to verify counts, investigate cleaning decisions, recover from interrupted work, and avoid treating a transformed spreadsheet as the original source.

Does an autocomplete suggestion represent search volume?

No. A suggestion can show that a phrase appears in an autocomplete footprint, but it does not establish monthly searches, traffic potential, ranking difficulty, or business value. Those are different claims requiring different evidence.

Should a zero-result keyword be discarded?

Not automatically. A zero-result seed may reflect phrasing, ambiguity, or limited provider coverage. A topic can still be valuable when first-party authority, audience need, portfolio fit, or other evidence supports it.

How should keyword research improve after publication?

Later research should compare the existing backlog with observed query and page evidence. Search Console can show search-performance patterns, while on-site behavior belongs to analytics. Neither source should be used to invent conclusions before a meaningful observation window exists.

See the wider content system

The AI-assisted SEO content pipeline article explains how this research layer connects to evidence planning, bilingual editorial production, website implementation, publication safety, and later observation.

Share

Found this useful? Share it.

LinkedInXCopy link

Discuss

Want to discuss the idea?

Discuss with me

Author

About the author

Huiyang Xie is a marketing professional based in Greater Vancouver, Canada, working across digital marketing, content, websites, AI-assisted workflows, and cross-cultural marketing. Her work explores how AI can support practical marketing processes while keeping factual review, brand judgment, and human approval in the loop.

Continue reading

Related work and reading

Reading

Building an AI-Assisted SEO Content Pipeline From Scratch

Open

Reading

How to Define “Done” in an AI-Assisted Workflow

Open