WP Internal Links
WordPress plugin: maps your internal links as a graph and finds the ones you haven't made yet. Works with WordPress's built-in AI connectors, any OpenAI-compatible endpoint, or no API key at all.
by Sanjay Shankar M · github.com/sanjuacodez/wp-internal-links
Install
No release zip yet. The repository archive installs, but the folder name will carry the branch suffix and updates will not flow:
wp plugin install https://github.com/sanjuacodez/wp-internal-links/archive/refs/heads/main.zipMaps how your WordPress content links to itself, and finds the internal links you have not made yet. Works with any AI provider you point it at — and still works without one.

The concept
Internal linking tools usually fail in one of two ways. Keyword tools match strings and suggest linking "SEO" on every page that says SEO. AI tools ask a model "which of these pages are related?" and hand back a list of pairs you then have to go find a place for by hand.
This plugin treats an internal link as something that has to be true in two different ways at once:
- The pages are about the same thing. Measured semantically, not by string overlap — so "market simulation" and "we ran 19 agents against our product" register as related even with no shared keyword.
- There is somewhere to put the link. A phrase describing the destination has to already exist in the source page's prose, in a spot where a link can legally go — not inside an existing link, a heading, a code block or a shortcode.
Either signal alone produces noise. Requiring both is what turns "these pages are related" into "put this link in this sentence" — a row you can action in one click instead of a research task.
The graph side is the other half: before you add links, you usually want to see which pages nothing points at. That view is built from the links that literally exist in your HTML — measured, not predicted.
How it works
Rebuild map & find links
│
├─ index ─ render each post → plain text, outbound links, keyphrases, embedding
│ (skips posts whose content hash has not changed)
├─ graph ─ invert outbound links into edges; count inbound; flag orphans
└─ analyze ─ similarity → candidate pairs → anchor must exist in linkable prose
→ score → store
Each stage runs in small browser-driven batches, so a large site never hits a PHP
timeout. Per-post data lives in postmeta (_ail_vector, _ail_phrases,
_ail_links_out, …), keyed by a content hash — edit a post and only that post is
re-indexed on the next run.
Scoring. 0.65 × similarity + 0.35 × anchor quality, plus +0.08 when the
destination has no inbound links at all. Similarity is calibrated per matching
mode (embeddings cluster in 0.5–0.9, TF-IDF in 0.05–0.3) and blended 60/40 with
how the pair ranks against the source page's own best candidate — so a short page
with a small vocabulary still gets a usable ranking.
Anchor quality rewards longer phrases over single words, rewards a phrase that is literally the destination's title, and rejects phrases that appear more than 8 times on the page or that appear on nearly every page in the corpus.
What it costs
The design point is that the expensive thing — asking a model to compare pages — never happens. Embeddings do candidate generation at roughly 1/1000th the price, and the chat model is a per-row tool you reach for, not a batch job.
Measured on a real 36-post site: 335,551 characters of body text, of which 182,662 characters (~45,700 tokens) are sent to the embeddings endpoint — each post is capped at 6,000 characters, since a post's topic lives in its opening.
Indexing (embeddings)
| Model | This 36-post site | A 500-post site |
|---|---|---|
text-embedding-3-small |
$0.0009 | $0.013 |
text-embedding-3-large |
$0.0059 | $0.083 |
A full re-index of a 500-page site costs about one cent. And you rarely pay even that: unchanged posts are skipped, so a normal re-run after editing one post re-embeds one post (~1,500 tokens, about $0.00003).
Ask AI (chat, per row, only when you press the button)
Measured request: ~448 tokens in, max_tokens 400 out, typical reply ~80 tokens.
| Model | Per row (typical) | Per row (worst case) | 100 rows |
|---|---|---|---|
gpt-5-nano |
$0.00005 | $0.00018 | $0.01 – $0.02 |
gpt-4o-mini |
$0.00012 | $0.00031 | $0.01 – $0.03 |
claude-haiku-4-5 |
$0.00085 | $0.00245 | $0.08 – $0.24 |
claude-sonnet-5 |
$0.00170 | $0.00490 | $0.17 – $0.49 |
And with no API key at all: $0.00
Leave the embedding model blank and matching falls back to TF-IDF cosine over your own corpus. The map, the orphan detection, the opportunity table, the anchor matching and the one-click insertion all work unchanged — only the similarity signal gets weaker. The 36-post site above ran end to end in 8 seconds for nothing, producing 96 mapped links and 88 opportunities.
Prices checked September 2026 (OpenAI) and against Anthropic's June 2026 rate card. Rates move — the token counts above are the stable part; multiply them by your provider's current rate.
Connecting an AI
There are two ways in, and the first one may already be done.
Option A — use WordPress's own AI connectors
WordPress bundles an AI Client in core (wp-includes/php-ai-client, verified on
7.0; the provider plugins declare 6.9+). Provider plugins — AI Provider for
OpenAI / Anthropic / Google, or the AI feature plugin — register into it
and store the credentials. If you have already set one of those up, pick the
preset WordPress AI (site's configured provider) and you are done: no key,
no base URL, no model name. The plugin reads the registered providers, shows you
which ones actually have credentials, and uses the site's own.
Settings → Internal Links → Settings → Provider: WordPress AI
You can pin it to a specific provider (openai / anthropic / google) or let
WordPress choose. Errors from the provider are passed through verbatim rather
than flattened into "request failed", because they are usually the actionable
part — a bad key, a model you lack access to, or an empty balance.
One caveat, and it matters: the AI Client exposes text, image, speech and
video generation — but no embeddings endpoint. EMBEDDING_GENERATION exists
in its capability enum with nothing behind it yet. So Option A powers the Ask
AI button, and embeddings are configured separately (below). Until that lands
in core, "WordPress AI only" means semantic matching falls back to TF-IDF.
Option B — point it at an endpoint directly
| Field | What it does |
|---|---|
| Provider | Preset. Custom is the escape hatch for anything else. |
| API base URL | Everything is appended to this. No trailing slash. |
| API key | Sent as Authorization: Bearer (or x-api-key for Anthropic). Blank leaves the stored key untouched. |
| Chat model | Used only when you press Ask AI on a row. |
| Endpoint paths | /chat/completions and /embeddings by default — change for gateways that mount them elsewhere. |
Embeddings are configured on their own
Because of the caveat above, the Embeddings section has its own base URL, key and model, each falling back to the chat connection when left blank. That means combinations like chat through WordPress's Google connector, embeddings through OpenAI work, and so does chat through Anthropic, embeddings through a local Ollama. Leave the embedding model blank to run on TF-IDF and spend nothing.
Press Test connection after saving — it does a real round trip and reports which path each half took, the model's reply, and the embedding dimensions it got back.
The wire formats
Two request shapes are implemented, which between them cover nearly everything:
- OpenAI-compatible —
POST {base}/chat/completionswith{model, messages, max_tokens}, andPOST {base}/embeddingswith{model, input: [...]}. This is the lingua franca: OpenAI, Azure OpenAI, OpenRouter, Together, Groq, Fireworks, DeepInfra, vLLM, LM Studio, Ollama's compatibility endpoint, LiteLLM and most in-house gateways all speak it. - Anthropic —
POST {base}/messageswithsystemhoisted out ofmessagesand theanthropic-versionheader.
If your provider speaks the OpenAI shape, pick Custom, give it your base URL, and it will work without a code change. If it needs an extra header or a tweaked payload, that is one filter:
add_filter( 'ail_request_args', function ( $args, $url, $settings ) {
$args['headers']['X-Org-Id'] = 'acme';
return $args;
}, 10, 3 );
Which models can connect
For chat (the Ask AI button) — any instruction-following model that returns
JSON on request. The task is small and well-specified (judge one link, pick an
anchor from a given passage), so cheap models do it well; there is no reason to
reach for a frontier model here. Verified shapes: OpenAI gpt-*, Anthropic
claude-*, and any OpenAI-compatible endpoint serving open models (Llama,
Qwen, Mistral, DeepSeek …). Non-JSON replies are handled — fenced ```json blocks
are stripped before parsing, and a malformed reply surfaces as an error on the
row rather than corrupting anything.
For embeddings — any endpoint implementing POST /embeddings with
{model, input: [...]} returning data[].embedding. Vector length does not
matter; the plugin reads whatever dimension comes back, and reorders by the
response's index field for providers that return rows out of order. Local
models work well here and cost nothing to run: nomic-embed-text or
mxbai-embed-large through Ollama or LM Studio are a good fit, since embedding
quality matters far more than model size for this task.
Two constraints worth knowing before you pick:
- Anthropic has no embeddings endpoint, and neither does WordPress's AI Client. Either one powers Ask AI fine; point the Embeddings section at a provider that has one (or leave it blank for TF-IDF).
- Embeddings are not interchangeable between models. Vectors from one model cannot be compared with vectors from another.
Changing the embedding model invalidates the cached vectors — they are not comparable across models. Tick Re-index everything on the next run after you switch.
Screens

- Link Map — force-directed graph of the real
<a>links between your posts. Node size is inbound links, a red ring means orphan. Click for a breakdown of what links in and out; drag to move, scroll to zoom, double-click to open. The layout is deterministic, so the map looks the same on every visit. - Opportunities — ranked table of source → destination with the anchor text highlighted in its surrounding sentence. Insert link writes the link into the post (creating a revision, so it is undoable from the editor), Ask AI refines it, Dismiss hides it from future runs.
- Settings — as above, plus post types, excluded IDs, score threshold and how many suggestions to keep per page.
Guardrails
- The model's suggested anchor is adopted only if that exact string is found in the source content. It cannot write new text into your posts.
- Insertion re-locates the anchor at apply time and uses
substr_replace()at a byte offset — your content is never regex-rewritten. wp_update_post()withwp_slash(), so a revision is created and backslashes in your content survive.- The API key is never echoed back into the settings form.
Extending
add_filter( 'ail_index_queue', $ids ); // what gets indexed
add_filter( 'ail_rendered_content', $html, $post ); // what the indexer reads
add_filter( 'ail_similarity_scale', array( $floor, $ceil ) ); // scoring calibration
add_filter( 'ail_stopwords', $words ); // keyphrase extraction
add_filter( 'ail_batch_size', $size, $phase ); // posts per AJAX step
add_filter( 'ail_request_args', $args, $url, $settings ); // extra headers / payload
add_filter( 'ail_link_markup', $html, $anchor, $target, $source ); // inserted markup
Known limits
- Pages built from custom blocks that store their text in block attributes (JSON inside the `` comment) produce few suggestions: there is no HTML body text to wrap, so no anchor can be inserted safely. Measured on this site: 60–90% of a normal post is linkable prose, versus 2–25% of a custom-block page. Those pages still appear in the map.
- The graph JSON is inlined on the map screen. Past a few thousand posts it should move to a REST endpoint with server-side filtering.
- Pair scoring is O(n²) in PHP. Fine into the low thousands of posts; beyond that it wants an approximate-nearest-neighbour index.
- WordPress's AI Client has no embeddings API yet, so the semantic half always needs a direct endpoint (or runs on TF-IDF).