WP Manifestindependent plugin directory
manifest / seo / wp-internal-links

WP Internal Links

WordPress plugin: maps your internal links as a graph and finds the ones you haven't made yet. Works with WordPress's built-in AI connectors, any OpenAI-compatible endpoint, or no API key at all.

by Sanjay Shankar M · github.com/sanjuacodez/wp-internal-links

★ 0stars
0forks

Install

No release zip yet. The repository archive installs, but the folder name will carry the branch suffix and updates will not flow:

wp plugin install https://github.com/sanjuacodez/wp-internal-links/archive/refs/heads/main.zip

Maps how your WordPress content links to itself, and finds the internal links you have not made yet. Works with any AI provider you point it at — and still works without one.

The link map: nodes are posts, edges are the internal links that actually exist. Selecting a node isolates its neighbourhood and lists what points at it.


The concept

Internal linking tools usually fail in one of two ways. Keyword tools match strings and suggest linking "SEO" on every page that says SEO. AI tools ask a model "which of these pages are related?" and hand back a list of pairs you then have to go find a place for by hand.

This plugin treats an internal link as something that has to be true in two different ways at once:

  1. The pages are about the same thing. Measured semantically, not by string overlap — so "market simulation" and "we ran 19 agents against our product" register as related even with no shared keyword.
  2. There is somewhere to put the link. A phrase describing the destination has to already exist in the source page's prose, in a spot where a link can legally go — not inside an existing link, a heading, a code block or a shortcode.

Either signal alone produces noise. Requiring both is what turns "these pages are related" into "put this link in this sentence" — a row you can action in one click instead of a research task.

The graph side is the other half: before you add links, you usually want to see which pages nothing points at. That view is built from the links that literally exist in your HTML — measured, not predicted.


How it works

Rebuild map & find links
   │
   ├─ index    ─ render each post → plain text, outbound links, keyphrases, embedding
   │             (skips posts whose content hash has not changed)
   ├─ graph    ─ invert outbound links into edges; count inbound; flag orphans
   └─ analyze  ─ similarity → candidate pairs → anchor must exist in linkable prose
                 → score → store

Each stage runs in small browser-driven batches, so a large site never hits a PHP timeout. Per-post data lives in postmeta (_ail_vector, _ail_phrases, _ail_links_out, …), keyed by a content hash — edit a post and only that post is re-indexed on the next run.

Scoring. 0.65 × similarity + 0.35 × anchor quality, plus +0.08 when the destination has no inbound links at all. Similarity is calibrated per matching mode (embeddings cluster in 0.5–0.9, TF-IDF in 0.05–0.3) and blended 60/40 with how the pair ranks against the source page's own best candidate — so a short page with a small vocabulary still gets a usable ranking.

Anchor quality rewards longer phrases over single words, rewards a phrase that is literally the destination's title, and rejects phrases that appear more than 8 times on the page or that appear on nearly every page in the corpus.


What it costs

The design point is that the expensive thing — asking a model to compare pages — never happens. Embeddings do candidate generation at roughly 1/1000th the price, and the chat model is a per-row tool you reach for, not a batch job.

Measured on a real 36-post site: 335,551 characters of body text, of which 182,662 characters (~45,700 tokens) are sent to the embeddings endpoint — each post is capped at 6,000 characters, since a post's topic lives in its opening.

Indexing (embeddings)

Model This 36-post site A 500-post site
text-embedding-3-small $0.0009 $0.013
text-embedding-3-large $0.0059 $0.083

A full re-index of a 500-page site costs about one cent. And you rarely pay even that: unchanged posts are skipped, so a normal re-run after editing one post re-embeds one post (~1,500 tokens, about $0.00003).

Ask AI (chat, per row, only when you press the button)

Measured request: ~448 tokens in, max_tokens 400 out, typical reply ~80 tokens.

Model Per row (typical) Per row (worst case) 100 rows
gpt-5-nano $0.00005 $0.00018 $0.01 – $0.02
gpt-4o-mini $0.00012 $0.00031 $0.01 – $0.03
claude-haiku-4-5 $0.00085 $0.00245 $0.08 – $0.24
claude-sonnet-5 $0.00170 $0.00490 $0.17 – $0.49

And with no API key at all: $0.00

Leave the embedding model blank and matching falls back to TF-IDF cosine over your own corpus. The map, the orphan detection, the opportunity table, the anchor matching and the one-click insertion all work unchanged — only the similarity signal gets weaker. The 36-post site above ran end to end in 8 seconds for nothing, producing 96 mapped links and 88 opportunities.

Prices checked September 2026 (OpenAI) and against Anthropic's June 2026 rate card. Rates move — the token counts above are the stable part; multiply them by your provider's current rate.


Connecting an AI

There are two ways in, and the first one may already be done.

Option A — use WordPress's own AI connectors

WordPress bundles an AI Client in core (wp-includes/php-ai-client, verified on 7.0; the provider plugins declare 6.9+). Provider plugins — AI Provider for OpenAI / Anthropic / Google, or the AI feature plugin — register into it and store the credentials. If you have already set one of those up, pick the preset WordPress AI (site's configured provider) and you are done: no key, no base URL, no model name. The plugin reads the registered providers, shows you which ones actually have credentials, and uses the site's own.

Settings → Internal Links → Settings → Provider: WordPress AI

You can pin it to a specific provider (openai / anthropic / google) or let WordPress choose. Errors from the provider are passed through verbatim rather than flattened into "request failed", because they are usually the actionable part — a bad key, a model you lack access to, or an empty balance.

One caveat, and it matters: the AI Client exposes text, image, speech and video generation — but no embeddings endpoint. EMBEDDING_GENERATION exists in its capability enum with nothing behind it yet. So Option A powers the Ask AI button, and embeddings are configured separately (below). Until that lands in core, "WordPress AI only" means semantic matching falls back to TF-IDF.

Option B — point it at an endpoint directly

Field What it does
Provider Preset. Custom is the escape hatch for anything else.
API base URL Everything is appended to this. No trailing slash.
API key Sent as Authorization: Bearer (or x-api-key for Anthropic). Blank leaves the stored key untouched.
Chat model Used only when you press Ask AI on a row.
Endpoint paths /chat/completions and /embeddings by default — change for gateways that mount them elsewhere.

Embeddings are configured on their own

Because of the caveat above, the Embeddings section has its own base URL, key and model, each falling back to the chat connection when left blank. That means combinations like chat through WordPress's Google connector, embeddings through OpenAI work, and so does chat through Anthropic, embeddings through a local Ollama. Leave the embedding model blank to run on TF-IDF and spend nothing.

Press Test connection after saving — it does a real round trip and reports which path each half took, the model's reply, and the embedding dimensions it got back.

The wire formats

Two request shapes are implemented, which between them cover nearly everything:

  • OpenAI-compatible — POST {base}/chat/completions with {model, messages, max_tokens}, and POST {base}/embeddings with {model, input: [...]}. This is the lingua franca: OpenAI, Azure OpenAI, OpenRouter, Together, Groq, Fireworks, DeepInfra, vLLM, LM Studio, Ollama's compatibility endpoint, LiteLLM and most in-house gateways all speak it.
  • Anthropic — POST {base}/messages with system hoisted out of messages and the anthropic-version header.

If your provider speaks the OpenAI shape, pick Custom, give it your base URL, and it will work without a code change. If it needs an extra header or a tweaked payload, that is one filter:

add_filter( 'ail_request_args', function ( $args, $url, $settings ) {
    $args['headers']['X-Org-Id'] = 'acme';
    return $args;
}, 10, 3 );

Which models can connect

For chat (the Ask AI button) — any instruction-following model that returns JSON on request. The task is small and well-specified (judge one link, pick an anchor from a given passage), so cheap models do it well; there is no reason to reach for a frontier model here. Verified shapes: OpenAI gpt-*, Anthropic claude-*, and any OpenAI-compatible endpoint serving open models (Llama, Qwen, Mistral, DeepSeek …). Non-JSON replies are handled — fenced ```json blocks are stripped before parsing, and a malformed reply surfaces as an error on the row rather than corrupting anything.

For embeddings — any endpoint implementing POST /embeddings with {model, input: [...]} returning data[].embedding. Vector length does not matter; the plugin reads whatever dimension comes back, and reorders by the response's index field for providers that return rows out of order. Local models work well here and cost nothing to run: nomic-embed-text or mxbai-embed-large through Ollama or LM Studio are a good fit, since embedding quality matters far more than model size for this task.

Two constraints worth knowing before you pick:

  1. Anthropic has no embeddings endpoint, and neither does WordPress's AI Client. Either one powers Ask AI fine; point the Embeddings section at a provider that has one (or leave it blank for TF-IDF).
  2. Embeddings are not interchangeable between models. Vectors from one model cannot be compared with vectors from another.

Changing the embedding model invalidates the cached vectors — they are not comparable across models. Tick Re-index everything on the next run after you switch.


Screens

The opportunities table: every row is a destination, the anchor text, and the sentence it already sits in

  • Link Map — force-directed graph of the real <a> links between your posts. Node size is inbound links, a red ring means orphan. Click for a breakdown of what links in and out; drag to move, scroll to zoom, double-click to open. The layout is deterministic, so the map looks the same on every visit.
  • Opportunities — ranked table of source → destination with the anchor text highlighted in its surrounding sentence. Insert link writes the link into the post (creating a revision, so it is undoable from the editor), Ask AI refines it, Dismiss hides it from future runs.
  • Settings — as above, plus post types, excluded IDs, score threshold and how many suggestions to keep per page.

Guardrails

  • The model's suggested anchor is adopted only if that exact string is found in the source content. It cannot write new text into your posts.
  • Insertion re-locates the anchor at apply time and uses substr_replace() at a byte offset — your content is never regex-rewritten.
  • wp_update_post() with wp_slash(), so a revision is created and backslashes in your content survive.
  • The API key is never echoed back into the settings form.

Extending

add_filter( 'ail_index_queue', $ids );                        // what gets indexed
add_filter( 'ail_rendered_content', $html, $post );            // what the indexer reads
add_filter( 'ail_similarity_scale', array( $floor, $ceil ) );  // scoring calibration
add_filter( 'ail_stopwords', $words );                         // keyphrase extraction
add_filter( 'ail_batch_size', $size, $phase );                 // posts per AJAX step
add_filter( 'ail_request_args', $args, $url, $settings );      // extra headers / payload
add_filter( 'ail_link_markup', $html, $anchor, $target, $source ); // inserted markup

Known limits

  • Pages built from custom blocks that store their text in block attributes (JSON inside the `` comment) produce few suggestions: there is no HTML body text to wrap, so no anchor can be inserted safely. Measured on this site: 60–90% of a normal post is linkable prose, versus 2–25% of a custom-block page. Those pages still appear in the map.
  • The graph JSON is inlined on the map screen. Past a few thousand posts it should move to a REST endpoint with server-side filtering.
  • Pair scoring is O(n²) in PHP. Fine into the low thousands of posts; beyond that it wants an approximate-nearest-neighbour index.
  • WordPress's AI Client has no embeddings API yet, so the semantic half always needs a direct endpoint (or runs on TF-IDF).